Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:19:33.933240Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 107 outbound references and 13 inbound Pith citation observations for arXiv:2411.18363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:19:33.933240Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.613268Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T02:04:26.359777Z
100 of 107 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c49a763-780b-45de-b3ad-724e2f0244e5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70a0432-8b30-45e4-a6bf-1d5a78d5c38d · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32836d25-12f2-4c15-bd69-5d42b3976bb6 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pixtral 12B
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a42f71-a392-45d9-b2b2-3cbaa805c1f7 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45a4f56-b556-43db-aefb-e2d4b3c3a0e6 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71826af8-7568-4bf1-bc9b-ab1544bfae08 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1b61d5-4e9c-4d97-8934-4b7736802cda · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Coyo-700m: Image-text pair dataset
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146bd734-fea4-463d-a617-9f530aa56289 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding End-to-end object detection with transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8970b8-d77c-4241-8f49-99cdb15e5af5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca999a6-3ed2-4824-b53a-430952dd752f · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68b7adb-325e-42fc-bc3f-ebd6a2762b75 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db34f812-49d6-4670-9d15-af0173f4aa64 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pix2seq: A Language Modeling Framework for Object Detection
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb280a1-fc0d-47e9-82fb-2092ce0e5a12 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pali: A jointly-scaled multilingual language-image model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a09cba-1bdc-489f-b76c-7b3de7cdc07a · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241ac1b8-5919-4888-9aad-0b3a69b7be11 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gonzalez, Ion Stoica, and Eric P
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f058612-92d3-489e-a933-6acc46081016 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6eefa9-d1b1-46b9-b45c-46c7e763d4ca · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding NVLM: Open Frontier-Class Multimodal LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f48ef1-a536-4061-a918-14a515830312 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afeebe99-5d8d-4ba4-ae3e-dca7a3a4dbfb · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Imagenet: A large-scale hierarchical im- age database
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b960f285-0db9-488b-8b22-0e20c266cc44 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7918d0a-81da-4f83-b4f1-b3a453d2efe6 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding An im- age is worth 16x16 words: Transformers for image recog- nition at scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb12e2d-fa7e-40de-a793-2c002d74c98b · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ca6b0e-b4db-4829-9ad5-b4df2e2e313b · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e50e920-6395-4c51-a79f-41d74f7b8f78 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89c20cc-76c0-4b01-8c00-e816cebdf990 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Lvis: A dataset for large vocabulary instance segmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbbcd26-de36-43a9-82ee-ad4614138064 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mask r-cnn
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ee50429-81f4-44da-bb38-b3047538be34 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Icdar2019 com- petition on scanned receipt ocr and information extraction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ad07f8-1f7b-4ddc-a91e-4d83bd6cd882 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b21873-11d7-4150-9eab-11a29c6f2f12 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding T-rex2: Towards generic object detec- tion via text-visual prompt synergy
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5d74d5-ab14-4d63-a8a2-bdb9adb071da · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a964daef-5fd3-44d6-9fe8-982e2ecf46a5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding A diagram is worth a dozen images
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fab3ea-1bdd-4494-a132-ab39c3ed0183 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Segment Anything
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce005b6-582b-4185-ba00-ebbf54d339f4 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding The open images dataset v4: Unified image classifica- tion, object detection, and visual relationship detection at scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd9a563-b0ef-4a01-992e-f118bbf77b27 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LISA: Reasoning Segmentation via Large Language Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a47fddb-585e-4bde-9989-2ed609e05e5f · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206cc295-fd15-4fb3-aad2-fc8a00480be8 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34022ed7-6c7b-4b9b-b58e-1d0a0da3565f · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802e4187-9296-4023-832b-b491efc50e33 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e63d1198-7f0c-4b6e-bc8d-e84d9693c4cc · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounded language-image pre-training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e376fe7e-29cf-4f28-b173-783ef072a483 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Evaluating object hallucination in large vision-language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80263683-1793-451d-b747-a8208f0cc706 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55873e8b-ff74-4978-b429-92c8535b0f80 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03250065-3e06-4edf-a99c-4269ecb23042 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245b8a8b-6018-45b7-89fb-3618ebecf326 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Vila: On pre-training for vi- sual language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea6f3b6-3f0c-4e31-8145-d806fae5eade · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f9e16a-ded3-4fcd-ae61-c5b315355dbf · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376387ea-5c5f-4575-87a3-8facbff5244d · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf46bb7-cd75-4769-b4ad-1f09d8a9e86e · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Improved Baselines with Visual Instruction Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefe776c-5636-4920-95f0-f98fa39dcada · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Visual instruction tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72ebb23-5906-4e13-b8c3-e6bbe368bb3b · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7e0f3e-3217-41eb-8e70-680987f1715e · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054af2fd-d678-4085-9848-0c951120dfd5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MMBench: Is Your Multi-modal Model an All-around Player?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ab74b2-b8dc-4fa1-b349-6a84fd32af55 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Ocrbench: on the hidden mys- tery of ocr in large multimodal models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7cd212-b136-426c-b6df-5f3e4774d97b · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Swin transformer: Hierarchical vision transformer using shifted windows
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51de555d-bcc4-4124-87cf-cc1909032f9a · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding A convnet for the 2020s
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e184c4b4-aed2-4705-968d-2d608eb88878 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Towards end-to-end unified scene text detection and layout analysis
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90033afd-268b-4609-a24f-234dfbff9486 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec05580b-55f9-47b3-8f33-4751a25e6efa · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding KOSMOS-2.5: A Multimodal Literate Model
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12308234-30e1-44ed-814e-28e2f36fb9b3 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4534563-8430-48b7-b1a8-c5a984c02213 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Generation and comprehension of unambiguous object descriptions
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 86f407a5-eaf7-469a-849a-5252c052f3ba · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6d93a9-769d-48be-b490-20f46a7cd755 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gpt-4v(ision) system card
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25cd984b-fce6-4e41-a3ee-6e93e8620074 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dadb70b8-5652-4eaa-97d5-19a2cab5ec6b · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Perceptiongpt: Effectively fusing visual perception into llm
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 695c8c57-982f-45f6-afe7-95b205ba2046 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Learning transferable visual models from natural language supervision
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 196b3d6d-c0c1-4c75-a44e-858f16d654ef · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Paco: Parts 11 and attributes of common objects
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f75b2626-cb3c-48ef-abd7-476173ae6985 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Glamm: Pixel grounding large multimodal model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3465c83b-5709-461b-82d7-7b104d721219 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Girshick, and Jian Sun
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e233440e-e0bd-4320-8839-c0e8c2d4740c · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23fd7e0-2519-4635-9a3c-4bd43e84d4b5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Generalized in- tersection over union: A metric and a loss for bounding box regression
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df2c1ad-5e8f-4522-af0d-e33ed65c0154 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b64f5af3-eca8-406c-b2d0-db2b54e7940a · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CrowdHuman: A Benchmark for Detecting Human in a Crowd
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e98316-3883-4cfc-9989-d427f362307c · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Objects365: A large-scale, high-quality dataset for object detection
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation df506477-f157-4d6d-9192-7b5a093d8d76 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db5c0a3-e366-42fc-a107-a04da1d506f9 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Towards vqa models that can read
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9e0afbd2-cbf9-4e2d-84d4-688df13d5798 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Hashimoto
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e0f781fb-9217-4ebc-a702-0b76aa3adebc · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29c3f7b-339a-4bcb-9a09-81316e169f4e · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Internlm: A multilingual language model with progressively enhanced capabilities
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51eaa3d-f744-476d-89b1-d1ffcf935919 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Cambrian-1: A fully open, vision-centric ex- ploration of multimodal llms
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation beef477f-9c6c-46d8-a5a7-a21083fb81e4 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c85656-9442-40bf-807e-5c882275d847 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf20ef6e-2ef1-4fba-bd1e-9ebc317301a6 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30464c58-6124-4646-ada6-e74e40fd90e5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585288b8-058f-4bc6-a95a-b5a3a7a589b7 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b495736b-61db-4173-a1ef-e5670de323a0 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dde6206-e697-4acf-959d-817503224947 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Florence-2: Advancing a unified representation for a va- riety of vision tasks
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 33b7389b-5513-4a1d-9481-c34d3bcb9e98 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24a736c-29db-4215-baa7-c81b097b72d5 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cead7cea-608a-445e-9452-62dd927f3976 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding xgen-mm (blip-3): A family of open large multimodal models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e92401f-df96-4bea-9e3c-f3172154ebf9 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen2.5 Technical Report
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b81d476-9c47-421c-a7eb-51cd12d2db36 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e09333-7e0d-484b-959b-c8401490223e · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70984f70-42c8-4e13-93e4-30c46108de17 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Modeling context in referring expres- sions
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76703c34-0d78-4813-a605-f8e98b935373 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24e2520-c16b-4e88-8344-794e34c4f095 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Osprey: Pixel understanding with visual instruction tuning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 70e6fbe0-9b32-4049-9a34-519dc19d1341 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f46053d-b3fa-41d8-94e2-129f3ee57a77 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding From recognition to cognition: Visual commonsense rea- soning
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ce95c1e9-0208-4384-a744-d0282c005580 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28a04c0-5913-4f92-8649-fb0a03a5f1a2 · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Griffon: Spelling out all object locations at any granularity with large language models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 781524ef-e52a-4a72-8e29-d2400b40ec8f · outbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66437ede-5786-4b06-80fd-7bf12bb9742e · inbound
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9de07c2-fa78-4dac-b2cb-26ffb269408d · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4fafc4-c40b-4f2c-abba-fae1b630d798 · inbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21bcd62c-d96f-4fc7-a957-29b10de5adb9 · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4bef3e-e2f9-4880-8b30-7aaf303fa134 · inbound
Grounding Everything in Tokens for Multimodal Large Language Models ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0b42aca-156c-4d4e-a81d-bc1744630647 · inbound
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91885753-a2b4-495b-87a1-4de3389b2bfd · inbound
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c321b45e-0f05-445c-9f31-6e0f27040570 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0ac54816-6a0b-4701-8b6b-f8cd2e3ab968 · inbound
HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd0898ff-794f-494b-91c6-da4a1dec1311 · inbound
Vision as Unified Multimodal Generation ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e0c9b431-5f79-46fc-9a18-833a08027fef · inbound
Foundation-Assisted Active Learning for Object Detection Annotation ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d972d2-c497-4c56-b402-548652603d6c · inbound
ReferTrack: Referring Then Tracking for Embodied Visual Tracking ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1f35da-7181-45c1-9d20-854025286865 · inbound
Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.