Pith. sign in

Paper Citation Record · LEDGER

LMM-Det: Make Large Multimodal Models Excel in Object Detection

As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.18300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18300 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:41.385870Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58ed4f22-b4f8-40ad-a4d9-359276a652bb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.331412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.111541Z digest=sha256:27d1ed75a023cb8264d3ee50b1bfaee95968c3639b897d8f046aa5174f68f146

Observation 27805ef5-c316-4dcb-8327-ef1aafee5b34 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.118189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.118189Z digest=sha256:ded1619aa588d935512602c10abe70048ee645e677831ebd6e5becd4c1e67450

Observation c03da974-5d73-4943-9c95-182c33f9a3a4 · outbound

This paper cites Cascade r-cnn: Delving into high quality object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Cascade r-cnn: Delving into high quality object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.316592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.124440Z digest=sha256:c15ba7a9bcc17239394540b33bf28bd280eb3d24a2d4bc86a90493731286ec8b

Observation d6888f5f-6120-4b14-bb59-2471112cb795 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Honeybee: Locality-enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.300835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.129824Z digest=sha256:096b47b26744177eb443b77b2f443182adbe239a04dc48b353754c30c7649431

Observation 9d2577a7-ffdb-4c06-a5e5-63e52e641ca9 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

LMM-Det: Make Large Multimodal Models Excel in Object Detection X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.134807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.134807Z digest=sha256:baaf65290d4ba59df9824c7071ac8f99e9b2e347b8859c4dfdd8b807199c8598

Observation 5d46851c-81b4-46c9-a20b-822e1986d547 · outbound

This paper cites Allava: Harnessing gpt4v- synthesized data for a lite vision-language model, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Allava: Harnessing gpt4v- synthesized data for a lite vision-language model, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.285513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.140378Z digest=sha256:f0032b8ad432a364c2bbda38b479ff4f59bcc9acc1b06f0ef3e74dd29783aa75

Observation 5d0fbbc1-6be6-4921-b7d9-c4527c2acd88 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.145863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.145863Z digest=sha256:d855c48db314c434494454bb87bc4c4e469fa2b7423fb4bfc242f8b5217d3135

Observation f2681c6b-841a-4a05-95db-0ff91859d98d · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.150411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.150411Z digest=sha256:2ad03b8f62c6802afad0573ef4619d7b46c8af27431058c6fca18cef309f91d2

Observation a80e0917-a2be-41b4-a030-a1668d3b4f21 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Glm: General language model pretraining with autoregressive blank infilling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.155155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.155155Z digest=sha256:89689a16be1c5ee83acb39445433a936bbfd242ea8ad3e0fd1d5f94f4c140043

Observation d39d1c4f-789e-4370-bc1d-855a369de980 · outbound

This paper cites LVIS: A dataset for large vocabulary instance segmentation.

LMM-Det: Make Large Multimodal Models Excel in Object Detection LVIS: A dataset for large vocabulary instance segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.159873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.159873Z digest=sha256:d3122bf3db9185074b79438999520b3fb4ce8e84108ac7781daade2d90470e84

Observation 344af88c-0b80-4af1-9c1c-219d6367009c · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.164283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.164283Z digest=sha256:2747afa6320634e3f833a30a610b627663dbdd21a61d3a0a293a2f79ba37c3bf

Observation 35d22764-04c0-456f-8ac3-6dd7db62aadc · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.169091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.169091Z digest=sha256:807b27054f80043c67ccfd4d548ec7087069f09ff5c9c38e287542be8cd7e6ee

Observation b80f73ab-759f-4a65-9d7a-ee91caaefc00 · outbound

This paper cites Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.249214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.173863Z digest=sha256:a1f7bdcbc2cd267eca3529550d0cdf60f8b8ba46fdf7600ca40d804ec86b2d2a

Observation 1929c7a6-2409-45b9-b92f-c3a2d481be66 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.232628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.178318Z digest=sha256:29b9a2b84830ab19624b5543b6b63ac743ac45dca6a93e743fea66b90d295220

Observation 51dc8c80-68d7-4889-8fcd-79fc9d9758a3 · outbound

This paper cites Detrs with hybrid matching.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs with hybrid matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.216107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.182911Z digest=sha256:d627cb4b16bd82bc55a410532346b7ee8ad214c71db115b6c05286cfe5388865

Observation 9a0e7103-7d10-4fe2-9e0e-4f07eb61bc5c · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.187182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.187182Z digest=sha256:c77655810c00366d47610a7425d267f69854220cdd6a120ebbeb2718c27f44bd

Observation 9800eeba-e1b2-4715-a7a4-646b4b1401ad · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Openimages: A public dataset for large-scale multi-label and multi-class image classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.187371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.192354Z digest=sha256:e58bf370736afa9137da54483a15bec34b5c4761c4e934880bb56fd3a79ccd9b

Observation d39e0ca0-c515-4ee9-844a-d43e7d6201cd · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Building and better understanding vision-language models: insights and future directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.196489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.196489Z digest=sha256:3834823f8a0400c48fd729d3299cba53399eb59022ce54a45999f82c6cfad5c0

Observation 476c6df1-cf32-4b2d-aa5b-90eddafff098 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.171495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.201411Z digest=sha256:3fa0b0c9f6887cbce6d2bb4a012de681c38ebf154777b21af1f48f239844967f

Observation 565cd5ab-fd74-4391-99f2-b2b06f98ee1e · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Monkey: Image resolution and text label are important things for large multi-modal models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.154303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.205867Z digest=sha256:eee36c3f6a4995ba5292eb9ecd1db3fe1817b26144b9329f42db4ff961a19235

Observation 98281367-7eab-4f27-84a6-8d5c990339cc · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Generative region-language pretraining for open-ended object detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.137555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.210434Z digest=sha256:270112613b92fa329f91b8e38665271f193186313385838af66d0d5d481b95c0

Observation 082175c0-5c53-4bd6-8a3e-494d025f87be · outbound

This paper cites Microsoft coco: Common objects in context.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.121477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.214842Z digest=sha256:fe5cef1b823fea73ec1b44646b4ac665b7c2a47b8ae10d934f8270b8b8323758

Observation ee13eab0-2b8b-45fe-a68a-0c7b0b986ce9 · outbound

This paper cites Visual instruction tuning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.103453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.219573Z digest=sha256:570d271b340d171cdaf65db86a8a23fda8246afc4346159de97cc0ddab378832

Observation 1b120a8b-eab7-47e7-bd7f-7b30a3a458b1 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.088137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.224159Z digest=sha256:637a3e4218aeb333af1da582f7ab67c8bf4a17087dba22d4ed9dd20b9ba18271

Observation e431bbaa-d72d-4f2a-8b53-69f2174f4a63 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.229099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.229099Z digest=sha256:084ec003261e5eaf858f4d7d0441a2cfdad42fb71bcde8ad1c55f9a4d2440163

Observation 98756f12-295d-4452-a514-451e1932840d · outbound

This paper cites Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:20:41.658332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.233757Z digest=sha256:cc0bde2f6bc237986674cd464b5134a3d430806bdd838180113f52d76fc75aa5

Observation c61919d6-a94b-425b-b8c5-06915cb95c05 · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.238757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.238757Z digest=sha256:a5ecd6c725a13bcb0ae93725707cde4e8ea04756a5749ed772af1bb650cc5720

Observation 052f5567-27c8-4b4c-be71-d7197552d3df · outbound

This paper cites Scal- ing open-vocabulary object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Scal- ing open-vocabulary object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.071999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.243496Z digest=sha256:3e0bbbfe333e6ac57468e484db221e80e7440aa8b97f3f27b69fbd9b0b200307

Observation a00781fd-abed-48a2-acaa-09997f4b0545 · outbound

This paper cites an unresolved cited work.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:42.054615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.248336Z digest=sha256:f77dcde6679808be8688cc5806426fd58f55819147ad55d63e6299558da1f520

Observation 2f254c71-48af-45ae-932f-27a1aa2fa4ad · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.253210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.253210Z digest=sha256:daff4295425cc82c46040535d062a853a8c1fe41d80237613d321c5b73bfb527

Observation b5139c15-b74f-43f1-bee1-28ef7c17ab29 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.039031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.258353Z digest=sha256:e052b887848c93ff788d3bb990b373ca568d61de56c59add2f23ff2d00ffa1d0

Observation 0b4a0aea-8d96-4d95-9530-be0dd34651e0 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.262855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.262855Z digest=sha256:f44c624e1cc8d3c74932d3955e2a533cfedf581df3ad2b77d87cd72a9e5f88d1

Observation 837f13fa-8959-442a-b801-bc624c644b0d · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:42.010266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.267864Z digest=sha256:e20d61613bf6fd6a63d5b57e0473e7baded1f82956838826aa70359cccc5bf8e

Observation 6cd99452-18f4-4f55-99e5-0da38c2b3aa0 · outbound

This paper cites Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.993470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.272251Z digest=sha256:c73bd1169d6efbb8a5feca54da36c94e5ab11bc1b2e588bfe3b987eada5d8106

Observation 49ccb3ad-52a0-4065-a8cd-990bbf6c9447 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.978046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.276858Z digest=sha256:a6d0ba9dd5d2737f859e4921d9c728f1ce2178ca6a849a4a904e2f8c2fa85bbf

Observation b870a08f-6937-4ea0-895f-4d4287904f6e · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.962458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.281336Z digest=sha256:4c29b6656c0753e0b3654c42ffa05068134df21b35973f66f69ff796d53bf92f

Observation 0322734e-0a1e-4d52-b725-2d6e3fae8019 · outbound

This paper cites Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.947188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.285832Z digest=sha256:29ad2b53da9b578723d8df06f3bb28f1fdade74d2e5e0b032e5ac3499df720ab

Observation 468d53aa-7fe4-4c97-8da8-343b9e703d7b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.290110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.290110Z digest=sha256:3aead8f2566903bb3411542215df92b208cc65a577cc38b617bfdd97c8f5f8cd

Observation 8fb52397-f1fb-4d66-8565-f7bb61048fbf · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.295610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.295610Z digest=sha256:8e83820d4d31da6d4878f4ba10c19403707f78a7fbdba901e818e422a87ede60

Observation e79f9c03-8f8e-4163-8f48-d84cf427e9cc · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

LMM-Det: Make Large Multimodal Models Excel in Object Detection General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.300411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.300411Z digest=sha256:f0eb0214d65a5246838251c8c7bb3c67868c81b11601f2320c2e3d89912ad04b

Observation be62617d-1ea0-40d5-a7ab-2180920b719e · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Skywork: A More Open Bilingual Foundation Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.304607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.304607Z digest=sha256:10cf6ba13f072c00aac91613bf74194e5d9d4741e6567b28a47481c7b0e61cf6

Observation ff4647ae-7bcd-46f1-b7d6-d1feb6494acd · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.931671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.309479Z digest=sha256:79cc48ea2c9eb4e37abb69393667260f9e4cedd473ca4f5158e6c4aea3b94dc1

Observation 8ca45a51-e777-490b-9d47-ac8345db92a6 · outbound

This paper cites Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.915560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.315064Z digest=sha256:0a6f1566d0aa85eb428167581018048dd91b8750c4f47e1703453583b1b27c4f

Observation 046407b3-8617-4bad-93dd-d6babea75009 · outbound

This paper cites Ccmb: A large-scale chinese cross- modal benchmark.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ccmb: A large-scale chinese cross- modal benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.899307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.319711Z digest=sha256:3806909534091bdf3de6915d89d60bee82977a88d5abb60dda7e61d2e9399d22

Observation 4c9c66b6-a221-4452-863c-7827fa9e2d26 · outbound

This paper cites FG-CLIP: Fine-Grained Visual and Textual Alignment.

LMM-Det: Make Large Multimodal Models Excel in Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.324154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.324154Z digest=sha256:699861e6cdd5fffee3aa4f29688f4cb467b7f193eaf8dfcb453c15ac772ee68b

Observation a4af59aa-d814-470f-b77b-9845ee14c2f8 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Llava-cot: Let vision language models reason step-by-step, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.882457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.328936Z digest=sha256:ce3bc8dd123525bb8cd7b511c9abe8bd665ecec04e310d434b1b5dd44b8e6d58

Observation e7e450db-e393-40cb-b8f8-ba493cc0aafc · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.333364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.333364Z digest=sha256:6f23e602bf2d14983c7c6307e228eae353cf64ea4f0a9c5f0f97775efeda9d41

Observation 36d87560-cc15-4da0-9f18-62bc53828daa · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

LMM-Det: Make Large Multimodal Models Excel in Object Detection mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.337989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.337989Z digest=sha256:64c64d822a99b4511ddc761e6bb20829321eced3fc6aa26347aad5e257f3e051

Observation 25938f21-ddf7-4d62-b1a7-7c9a3aed47e2 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.343049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.343049Z digest=sha256:3976902dbe6d0a6efe1035e7b90e9944e42b69ec556ee9aa35124c40af0d7fed

Observation 5bd2b0d3-8bcc-4278-b687-13d9cf31877b · outbound

This paper cites Griffon v2: Advancing multimodal perception with high-resolution scaling and visual-language co-referring, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon v2: Advancing multimodal perception with high-resolution scaling and visual-language co-referring, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.865686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.347776Z digest=sha256:c4e5dffe7dc69b53ece22379784a47132e99d1e0aa8f7c5756dcd57ecfa79f8e

Observation 05456998-e84d-446b-87f2-f5cab2bb312f · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Griffon: Spelling out all object locations at any granularity with large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.849590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.352754Z digest=sha256:76a52715c8ca9caa53947e7a852c00b5cb94c41d4069069d523ad14a20d72a7c

Observation 5adb5b21-dda0-4a01-98fe-8b0fe5af8fa1 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.357690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.357690Z digest=sha256:02d3d35364899a50053686c1176281d1af3de3f7d47f75b19c345f257f75f813

Observation 8ff683b7-901f-48ef-a96f-1e6668a92687 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.362335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.362335Z digest=sha256:0de9dd8775ca86b7d964d49aefe20af3d53fb2d17875db49578d4cd120237a38

Observation acf63f30-37ca-40c5-8b35-65724dc9e9ef · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.367508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.367508Z digest=sha256:f4c8d241919dc41d97d6795f6a22cd942f901eee472a2dc75bdb222ba0099dd1

Observation acaece5a-8065-43db-ae5e-13efe8aae67f · outbound

This paper cites Detrs beat yolos on real-time object detection, 2024.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Detrs beat yolos on real-time object detection, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.833175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.372330Z digest=sha256:3473b80a5f1ba29db6fc38ff3beec81b723a66caafb63ae11f9381236846b09f

Observation 841aa40c-cb43-4908-b970-6244ff139a4d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LMM-Det: Make Large Multimodal Models Excel in Object Detection MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.377155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.377155Z digest=sha256:f41276aaa840254b451f4b9fc465cebb19b746d188d495cb65f386e70995ca02

Observation f1f79241-7ff3-4be4-a236-f2768dab37fe · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Deformable detr: Deformable transformers for end-to-end object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:41.816775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.381597Z digest=sha256:116b46d327a230bc3173a8c713846e37407f1a87ac23692733c43dab991d8ff5

Observation ae4fb308-6e4a-4e4b-92ce-a1acdf958aef · outbound

This paper cites Multi-step.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Multi-step

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:20:41.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:20:41.385870Z digest=sha256:8cda7a8901a42ef8a374fbc5b3edf37f01c9d9ee48fb09e3b36628aad6fbba83

Pith citing papers

No inbound Pith citation observations are available.