Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:41.385870Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.18300.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:41.385870Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 58ed4f22-b4f8-40ad-a4d9-359276a652bb · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 27805ef5-c316-4dcb-8327-ef1aafee5b34 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03da974-5d73-4943-9c95-182c33f9a3a4 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Cascade r-cnn: Delving into high quality object detection
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d6888f5f-6120-4b14-bb59-2471112cb795 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Honeybee: Locality-enhanced projector for multimodal llm
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9d2577a7-ffdb-4c06-a5e5-63e52e641ca9 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d46851c-81b4-46c9-a20b-822e1986d547 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Allava: Harnessing gpt4v- synthesized data for a lite vision-language model, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5d0fbbc1-6be6-4921-b7d9-c4527c2acd88 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2681c6b-841a-4a05-95db-0ff91859d98d · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80e0917-a2be-41b4-a030-a1668d3b4f21 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Glm: General language model pretraining with autoregressive blank infilling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39d1c4f-789e-4370-bc1d-855a369de980 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection LVIS: A dataset for large vocabulary instance segmentation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344af88c-0b80-4af1-9c1c-219d6367009c · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Efficient Multimodal Learning from Data-centric Perspective
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d22764-04c0-456f-8ac3-6dd7db62aadc · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection CogVLM2: Visual Language Models for Image and Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80f73ab-759f-4a65-9d7a-ee91caaefc00 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1929c7a6-2409-45b9-b92f-c3a2d481be66 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 51dc8c80-68d7-4889-8fcd-79fc9d9758a3 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs with hybrid matching
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9a0e7103-7d10-4fe2-9e0e-4f07eb61bc5c · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9800eeba-e1b2-4715-a7a4-646b4b1401ad · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Openimages: A public dataset for large-scale multi-label and multi-class image classification
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d39e0ca0-c515-4ee9-844a-d43e7d6201cd · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Building and better understanding vision-language models: insights and future directions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476c6df1-cf32-4b2d-aa5b-90eddafff098 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 565cd5ab-fd74-4391-99f2-b2b06f98ee1e · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Monkey: Image resolution and text label are important things for large multi-modal models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 98281367-7eab-4f27-84a6-8d5c990339cc · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Generative region-language pretraining for open-ended object detection
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 082175c0-5c53-4bd6-8a3e-494d025f87be · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Microsoft coco: Common objects in context
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee13eab0-2b8b-45fe-a68a-0c7b0b986ce9 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Visual instruction tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b120a8b-eab7-47e7-bd7f-7b30a3a458b1 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e431bbaa-d72d-4f2a-8b53-69f2174f4a63 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98756f12-295d-4452-a514-451e1932840d · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c61919d6-a94b-425b-b8c5-06915cb95c05 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 052f5567-27c8-4b4c-be71-d7197552d3df · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Scal- ing open-vocabulary object detection
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a00781fd-abed-48a2-acaa-09997f4b0545 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f254c71-48af-45ae-932f-27a1aa2fa4ad · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5139c15-b74f-43f1-bee1-28ef7c17ab29 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Learning transferable visual models from natural language supervision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0b4a0aea-8d96-4d95-9530-be0dd34651e0 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837f13fa-8959-442a-b801-bc624c644b0d · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6cd99452-18f4-4f55-99e5-0da38c2b3aa0 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 49ccb3ad-52a0-4065-a8cd-990bbf6c9447 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Objects365: A large-scale, high-quality dataset for object detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b870a08f-6937-4ea0-895f-4d4287904f6e · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0322734e-0a1e-4d52-b725-2d6e3fae8019 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 468d53aa-7fe4-4c97-8da8-343b9e703d7b · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb52397-f1fb-4d66-8565-f7bb61048fbf · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79f9c03-8f8e-4163-8f48-d84cf427e9cc · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be62617d-1ea0-40d5-a7ab-2180920b719e · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Skywork: A More Open Bilingual Foundation Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4647ae-7bcd-46f1-b7d6-d1feb6494acd · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8ca45a51-e777-490b-9d47-ac8345db92a6 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 046407b3-8617-4bad-93dd-d6babea75009 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Ccmb: A large-scale chinese cross- modal benchmark
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c9c66b6-a221-4452-863c-7827fa9e2d26 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4af59aa-d814-470f-b77b-9845ee14c2f8 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-cot: Let vision language models reason step-by-step, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e7e450db-e393-40cb-b8f8-ba493cc0aafc · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d87560-cc15-4da0-9f18-62bc53828daa · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25938f21-ddf7-4d62-b1a7-7c9a3aed47e2 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bd2b0d3-8bcc-4278-b687-13d9cf31877b · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon v2: Advancing multimodal perception with high-resolution scaling and visual-language co-referring, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05456998-e84d-446b-87f2-f5cab2bb312f · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon: Spelling out all object locations at any granularity with large language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5adb5b21-dda0-4a01-98fe-8b0fe5af8fa1 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff683b7-901f-48ef-a96f-1e6668a92687 · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf63f30-37ca-40c5-8b35-65724dc9e9ef · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acaece5a-8065-43db-ae5e-13efe8aae67f · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs beat yolos on real-time object detection, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 841aa40c-cb43-4908-b970-6244ff139a4d · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f79241-7ff3-4be4-a236-f2768dab37fe · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Deformable detr: Deformable transformers for end-to-end object detection
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ae4fb308-6e4a-4e4b-92ce-a1acdf958aef · outbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Multi-step
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.