Pith. sign in

Paper Citation Record · LEDGER

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

As of 24 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19857 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b995d4e-2191-4b15-b83f-b8fc2466296d · outbound

This paper cites The small-drone revolution is coming—scientists need to ensure it will be safe,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The small-drone revolution is coming—scientists need to ensure it will be safe,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.010000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.010000Z digest=sha256:216148a89e1985c145ab1847eb1f5b90915673fb5bbc945f777e5ee54ed02bce

Observation 49b07a8c-e707-4b7c-9a55-2c743b824f6b · outbound

This paper cites Champion-level drone racing using deep reinforcement learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Champion-level drone racing using deep reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.077507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.077507Z digest=sha256:610dc2fcd274f0e41e57c0a5a582c5139fed702d58100a0f8b38db2d89cb2729

Observation 761e62ab-1b1c-4192-b68f-2dad832a8038 · outbound

This paper cites Video object segmentation without tem- poral information,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Video object segmentation without tem- poral information,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.210959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.210959Z digest=sha256:4a4829bebe42a1f5c684e1d650fc6ec8a40b5a0893b350c10214fbe6f9ba41c4

Observation 5c39917d-b99c-406e-b8f3-61d24b26056c · outbound

This paper cites 3d question answering for city scene understanding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 3d question answering for city scene understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.290989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.290989Z digest=sha256:2828f2b5282a244edf6df3bafbf3a5dcd922686631d590523017eb4e4371f0ca

Observation 43efa3e9-286a-4e1b-be86-a3197c84754d · outbound

This paper cites Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.391352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.391352Z digest=sha256:0ae2bb620055069392088c0cb257724f19c039fb907fe8214b9572e364e5adcd

Observation 01ecd4ed-e87b-4b06-8dec-7ad48ad22ef2 · outbound

This paper cites Detecting flying objects using a single moving camera,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Detecting flying objects using a single moving camera,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.479024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.479024Z digest=sha256:32eff428754530c177ba47b13c232366bcb92242928bc0de7ae9ca2c450b0e3f

Observation 4f4d291c-5418-4571-8780-b59da9a6df49 · outbound

This paper cites Revisiting image-language networks for open-ended phrase detection,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Revisiting image-language networks for open-ended phrase detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.574070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.574070Z digest=sha256:eefdc57bbbdeb296514fdc99bebb775acbde27fe0e1b835b04eb1f727c444d37

Observation d2b93844-3363-4b7a-aff3-0037efbdcd01 · outbound

This paper cites Mevis: A large- scale benchmark for video segmentation with motion expressions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A large- scale benchmark for video segmentation with motion expressions,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.647764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.647764Z digest=sha256:c0b9e7055f7acc700fe5da585d70d5645ffc6ec6fc44e651e41545c8d140b64e

Observation d5431214-f946-448e-974c-5886b8e16458 · outbound

This paper cites Lamot: Language- guided multi-object tracking,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Lamot: Language- guided multi-object tracking,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.750581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.750581Z digest=sha256:899b11b0a715360fe17b746b41e2c5ef90207009eb8715107f5f487a725ad6ee

Observation a6822768-0774-409a-b3c4-642633a0a9f6 · outbound

This paper cites Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.793599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.793599Z digest=sha256:bde06e9c4b8e1bed62f98582b3c7293abfb9bae9c965c2dbb2bfceb0df84cd53

Observation 50c4b8f7-f820-4c1e-96ac-d23aae4cef2d · outbound

This paper cites Aerialmind: Towards referring multi-object tracking in uav scenarios,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Aerialmind: Towards referring multi-object tracking in uav scenarios,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.860865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.860865Z digest=sha256:c90f7a3144b5c59c4e19e2ec8d5a3403c2af75b1a73000a5a1f4cc9af3d21cef

Observation 6da0cf3b-dd9b-4492-acd3-8afc41de44ce · outbound

This paper cites Event-aware instructed assistant for referring video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Event-aware instructed assistant for referring video segmentation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.918717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.918717Z digest=sha256:30f4a5a50f470441f4485af5b1f28943d3a1110f89d02e37266380faeecffd27

Observation cd07ae9b-fc66-4d12-97e5-44f84799e2fb · outbound

This paper cites City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.000557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.000557Z digest=sha256:0fa171fb309f5602ac5c188c3503e77c0a6e4e5cc6621bdc53f06455551a3406

Observation 71600a21-e282-4709-8a33-42faab77c5fb · outbound

This paper cites Qwen2.5 technical report,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 technical report,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.085352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.085352Z digest=sha256:4d3886b32f4c4cc4667f763d0522ce6a212f2260e737c22523c40201c6f3e8f9

Observation 366e8549-1cbc-4afa-984b-b330d1f213b2 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.259681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.259681Z digest=sha256:03b0a04972ed20bae41e18031d770251eab16138ce20c457d9087722a5484748

Observation 305d387f-576c-47d0-99f8-484771a6ffed · outbound

This paper cites React: Streaming video analytics on the edge with asynchronous cloud support,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos React: Streaming video analytics on the edge with asynchronous cloud support,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.350558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.350558Z digest=sha256:3e94941f79022ce13c6ad73086ef06593cd77b50665d16d67e09fef7d9599f0e

Observation 5b29eab7-6a4d-4271-8344-f237d491cea0 · outbound

This paper cites Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.439765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.439765Z digest=sha256:182c02fe7e7e80a8e85802715c1b368d08836cc2b81ebb352e5367a200f02635

Observation 08fa439c-bdb6-4acf-8ed9-9f15f82ab321 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.515394Z digest=sha256:a6cb49465af54d2ef0d033e1c8ef3f4bb6218884a7109452f0e6317ac4ba6153

Observation b14fc791-78fa-4110-bcd4-66ae0c8a6cf8 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sharegpt4video: Improving video understanding and generation with better captions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.604426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.604426Z digest=sha256:5389d049cb2b6773698421d2b03c6ec9feaf8d1dab9d08fa70f4b008c7e34ac7

Observation 79b21db3-821c-4694-8687-61f39cbdae15 · outbound

This paper cites Mevis: A multi-modal dataset for referring motion expression video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A multi-modal dataset for referring motion expression video segmentation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.675678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.675678Z digest=sha256:16ce8734a1fed1c719c1c00e8ce95304075305188ad71499117941642b0f0892

Observation 844c35ca-a119-472a-8294-da0f114f5e6f · outbound

This paper cites Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.778744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.778744Z digest=sha256:57e8e869cf6bcdfa847c9d23cb4ee99a04bf2d20c58fa9bcae6e4be63845bd34

Observation 0e892caa-5ce7-4f56-bdb7-f50fa6c94cd4 · outbound

This paper cites Visa: Reasoning video object segmentation via large lan- guage models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visa: Reasoning video object segmentation via large lan- guage models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.829375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.829375Z digest=sha256:46c04d56c2291c86b495bef0324910a60029e9e5a44139250e460823519397cc

Observation eed37ea5-009d-4755-b182-9cc0b34b6b93 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.889990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.889990Z digest=sha256:3ca46dbf9caec0d2d131e5598c038d00d1213adb1bb69ba6d48424f6af20c11a

Observation 26c53c5e-2219-4323-a01b-152283af7035 · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.985730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.985730Z digest=sha256:e7ba8c689e1e8ff1894767619230468a8dd35ee0b4167ca043220855db1ca238

Observation 166f5840-37de-4732-b725-fca92a964b90 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.065979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.065979Z digest=sha256:32897df74774d489e99098fd5c8c72c5433c331905bf2a33c77f68b406ee2b59

Observation 4af657c0-610a-4099-a718-52299e96ebc4 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.135856Z digest=sha256:a00828881b85be482db2225126ceb64734d447b0ef4c799fa559fe57202605dd

Observation 7a0db379-3b2a-4df5-b8ca-796f9e2f9a06 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Glus: Global-local reasoning unified into a single large language model for video segmentation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.233507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.233507Z digest=sha256:9dcff51bcd4879c1fe73e73c91ff1cbe3b2bcaa6feef2a82c553a5c20c48b2af

Observation 5c1cf94f-b519-42a1-b153-da195082896f · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The devil is in temporal token: High quality video reasoning segmentation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.328278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.328278Z digest=sha256:7e0b295f225e03eded1a3e41230de2875c4e6d5ea43728c89f5d5c5c732d936e

Observation 1fc75169-9d8b-4957-8153-a284702ec0b5 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Instructseg: Unifying instructed visual segmentation with multi-modal large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.435658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.435658Z digest=sha256:b8c53eba02debc2154d3ffd625609c576d548be95579e1d10ac8f957d9e601e8

Observation df21c549-3a59-43ec-94da-f40429c71d43 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Geochat: Grounded large vision-language model for remote sensing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.532666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.532666Z digest=sha256:322e6f2a278096fcbc8f528f0185e1058bddfade5ce44d46b081d34c90ac69da

Observation aa2add77-efb9-4878-be89-fdcd1b629294 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.613930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.613930Z digest=sha256:d52e23e2bc2b82ca43d0dd68f8cf43d7b9777692612e62d49fff8ba66e5ac764

Observation 2ef00eab-20ef-407f-b9a5-a1423abd28ee · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.690981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.690981Z digest=sha256:8518e6d794b296789f80b08d2c61bc33e72293e1d60d8b5b603ffd9ec65e53a4

Observation 1b7777f8-6659-4ffb-9d4d-397c763d51cd · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Kimi K2.5: Visual Agentic Intelligence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.758425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.758425Z digest=sha256:f6714d6dcd3e63132553663b4dfb6a9aa1346ea05214f248beebbc8aa594fac4

Observation beedfb6f-dfe7-4f2b-9f22-bf52ce0350a9 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen3.6-Plus: Towards real world agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.831961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.831961Z digest=sha256:e0c07babb27159ce21704b1a1d8f34378d0e61da49d146c95b0e610efedfb296

Observation 211f0754-f661-4860-b3d8-3eb118d94675 · outbound

This paper cites Visual instruction tuning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.921225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.921225Z digest=sha256:e0c81a8079851cb35fe660df8db88756bd74b97f33cb85331c394a2ae899944a

Observation 49e040f5-7cb0-4289-9175-9dbffc0da9c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.990599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.990599Z digest=sha256:ff678bd91895f11c7964de8fb02759434d4d34e0c8f0e3538669c28716058732

Observation e7b2fe0a-3f59-41ca-9521-652d554e1f04 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.071908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.071908Z digest=sha256:c095b3cd727a36a88e84079481dd27265fc29d175d546519742bce4d5554b924

Observation 9f512ed1-8155-4f28-bd60-74bb222a0f54 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A fast and accurate one-stage approach to visual grounding,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.133296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.133296Z digest=sha256:1d528d1103b5a049fe695a00d6f707e481492c5d3d3ab89d3fc9440bb2486939

Observation dfd9d33c-78e5-40bc-b1bf-c29bba6295cf · outbound

This paper cites Improving one-stage visual grounding by recursive sub-query construction,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving one-stage visual grounding by recursive sub-query construction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.252921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.252921Z digest=sha256:0f310e86b62f5ed4c521af0f5a097e30eb80ad8b9af53b8ff1370ea167911166

Observation 994a86ff-27c6-4364-b2da-a43282072bfd · outbound

This paper cites Referring transformer: A one-step approach to multi-task visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Referring transformer: A one-step approach to multi-task visual grounding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.355563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.355563Z digest=sha256:9e3112a51ec9aa54974893c14f6c7d1d520b730a4fd2ec4f4447fd9f7cf5deac

Observation ed070ee0-2232-42c8-b209-4f4b6c51f151 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Transvg: End-to-end visual grounding with transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.438359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.438359Z digest=sha256:c6444fb42f23e6f611ac446cbf042b6484ff0795e487b0bb75407251cc34cb2a

Observation a3617427-5543-4c34-81cf-13884fbc05a6 · outbound

This paper cites Improving visual grounding with visual-linguistic verification and iterative reasoning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving visual grounding with visual-linguistic verification and iterative reasoning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.518896Z digest=sha256:6721f91bcadcedcd6cb3efe59f9546ba12d4c1e45991d2e6138218ddd1f1c568

Observation 15da3c5a-576d-4eba-b291-7e1adf5e5b66 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Seqtr: A simple yet universal network for visual grounding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.607850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.607850Z digest=sha256:5128fa3740fb1b85b68f20e31ae2ac7742ab112a8e8747396a14d52ab9261b1f

Observation 2c406c7c-c960-4212-b552-37e5b05d6790 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.682676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.682676Z digest=sha256:2708ad8c221e4682a3ba5d83575f1cd5e91bc52ec149405f04b2e3d99b25fd1a

Observation 8f6255a7-0317-45e2-98b1-3e35f30b41d7 · outbound

This paper cites A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.745993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.745993Z digest=sha256:f214eb8376b9fec76c1c37578ccf04f97c73cb2117201efae8506f1e983b56c6

Observation 2bc27e84-0bb9-4c6c-9f13-e9a704e31956 · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Polyformer: Referring image segmentation as sequential polygon generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.814543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.814543Z digest=sha256:9426d6364a60184a2f37143219df8e5fc73e584407aeb8b1fcb0b82337e05a3e

Observation c2c19032-b299-4504-a5f6-cad0810c8554 · outbound

This paper cites Qwen2.5 Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.188553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.188553Z digest=sha256:2fcf8874356fe3186f48f0482dce43328f8d9dbc3dadfc2fdee726e454144c47

Pith citing papers

No inbound Pith citation observations are available.