Pith. sign in

Paper Citation Record · LEDGER

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

As of 20 August 2026, this Paper Citation Record lists 100 of 107 outbound references and 13 inbound Pith citation observations for arXiv:2411.18363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18363 v3

Coverage vector

measured 100 of 107 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:19:33.933240Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.613268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.359777Z

Reference resolution

100 of 107 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c49a763-780b-45de-b3ad-724e2f0244e5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.395298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.395298Z digest=sha256:06f4958c2807a62b6ddb8e3c01789f81ed18fb4d96476b4fe3d9b50e7893218d

Observation a70a0432-8b30-45e4-a6bf-1d5a78d5c38d · outbound

This paper cites GPT-4 Technical Report.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.402612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.402612Z digest=sha256:d088fe81a4935a6c8feee8a481fdb5a5b6c4d0a59831cc0eeedbc182881d4a76

Observation 32836d25-12f2-4c15-bd69-5d42b3976bb6 · outbound

This paper cites Pixtral 12B.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pixtral 12B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.407841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.407841Z digest=sha256:b43db53bf5251f3c3499685e5bc7f1f0a0109b17910a469de2c922bef9ce9e84

Observation d5a42f71-a392-45d9-b2b2-3cbaa805c1f7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.413795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.413795Z digest=sha256:1f8a5428aa00751219927046bee4a77fb8ab6a37ad31ff72ba235ecdc7be5d25

Observation b45a4f56-b556-43db-aefb-e2d4b3c3a0e6 · outbound

This paper cites Qwen Technical Report.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.419113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.419113Z digest=sha256:8f1fcf3e0addbad8c61e86445398a61c88a7c95a759965a4fed491a11f0e89f0

Observation 71826af8-7568-4bf1-bc9b-ab1544bfae08 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.424647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.424647Z digest=sha256:a2642ef26396868ee739f147dcee8912db5a0218decc0f0badef999edb9a13f1

Observation cb1b61d5-4e9c-4d97-8934-4b7736802cda · outbound

This paper cites Coyo-700m: Image-text pair dataset.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Coyo-700m: Image-text pair dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.430385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.430385Z digest=sha256:ed02193b0c36bfe380baf50d147fa5733cbe47c2aa18e0fe3e9bb986d3187e61

Observation 146bd734-fea4-463d-a617-9f530aa56289 · outbound

This paper cites End-to-end object detection with transformers.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding End-to-end object detection with transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.435695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.435695Z digest=sha256:de1db35a355fb5f7ae02bc593a717bad1129b2346b94cabbffcba5b8e7f94873

Observation 7e8970b8-d77c-4241-8f49-99cdb15e5af5 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.440573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.440573Z digest=sha256:2a6c2ffbec357dd2cbc847ca65122ec77c48b4dd80e3373e4170d29d38d13a25

Observation cca999a6-3ed2-4824-b53a-430952dd752f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.445705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.445705Z digest=sha256:02dbdde93434bd53a5a32289a35e0562898cccc098177ca0c454b6233a0cd403

Observation d68b7adb-325e-42fc-bc3f-ebd6a2762b75 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.452132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.452132Z digest=sha256:76e76bb26b0cdad814c8f259e5575ce207e5b6a4b38102652e36540afad2b428

Observation db34f812-49d6-4670-9d15-af0173f4aa64 · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pix2seq: A Language Modeling Framework for Object Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.457024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.457024Z digest=sha256:2b3b7de473e405fa12823a74d38de5ef0799fad986e8e09d943ebafb9ce9a72b

Observation 2eb280a1-fc0d-47e9-82fb-2092ce0e5a12 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Pali: A jointly-scaled multilingual language-image model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.462248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.462248Z digest=sha256:07c5adfb7cbb4405a7c9825066db21e67d5246675b80aec9ed5577a985e618aa

Observation a1a09cba-1bdc-489f-b76c-7b3de7cdc07a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.466916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.466916Z digest=sha256:8b997ca96ce9c913c888ce140a2e9f3d335b46ee7d5224fe2060955b62a1f27b

Observation 241ac1b8-5919-4888-9aad-0b3a69b7be11 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.471677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.471677Z digest=sha256:fbea46b3c68899fbd32ff6a31017fa74e3b99f05b5e66db386093e7feac15cd0

Observation 6f058612-92d3-489e-a933-6acc46081016 · outbound

This paper cites an unresolved cited work.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.476539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.476539Z digest=sha256:9cf84c4d5b105c1e4bd70d8b8756b2817e71cf9ac6af82e610a7c2e1a06a414c

Observation 1f6eefa9-d1b1-46b9-b45c-46c7e763d4ca · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding NVLM: Open Frontier-Class Multimodal LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.481133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.481133Z digest=sha256:b0545617ac023464c73176716b970e6d0fdca2dfda4e5d309e8efb56b19decd4

Observation e7f48ef1-a536-4061-a918-14a515830312 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.487065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.487065Z digest=sha256:901af09723969f01b6ac407ee5a4d3bfb1a21356cd9e72828426a7a00ebab854

Observation afeebe99-5d8d-4ba4-ae3e-dca7a3a4dbfb · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Imagenet: A large-scale hierarchical im- age database

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.492658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.492658Z digest=sha256:37fed4a3a119c830c66b6e1d90e9f7a3d9b2ea582ddf0215fbfbe878a95d0aa4

Observation b960f285-0db9-488b-8b22-0e20c266cc44 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.497294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.497294Z digest=sha256:c4c518d04a1fe89bf1362d9a36531515afc0947a812ba1b756f26fd7b3122a54

Observation a7918d0a-81da-4f83-b4f1-b3a453d2efe6 · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.502100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.502100Z digest=sha256:32d5bcfa3b3d0c87bfb502e1531d74cdb788d5f9353a21b3bc6be5452b961356

Observation 6bb12e2d-fa7e-40de-a793-2c002d74c98b · outbound

This paper cites The Llama 3 Herd of Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.507002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.507002Z digest=sha256:625b636af8f745e39a840cfc4e1d818e37976947f853f283a8fe9f5922cb45f8

Observation 59ca6b0e-b4db-4829-9ad5-b4df2e2e313b · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.511859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.511859Z digest=sha256:0a8addcf309930b08d787d23181a2f81da4794f4b0c2ae9606018eedec19f371

Observation 0e50e920-6395-4c51-a79f-41d74f7b8f78 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.517203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.517203Z digest=sha256:066a9717b3d2ff0309e7a2ee5a9daa45a238f648e2590cc98a356a50f9194cde

Observation d89c20cc-76c0-4b01-8c00-e816cebdf990 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Lvis: A dataset for large vocabulary instance segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.522606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.522606Z digest=sha256:9418d53719ebcaf0ae5b1e757d333df0fcc07358f852636fc0c498f37fc82a0c

Observation 1dbbcd26-de36-43a9-82ee-ad4614138064 · outbound

This paper cites Mask r-cnn.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mask r-cnn

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.528216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.528216Z digest=sha256:f7c691d0ae5b4205c636c9194703095433efc3ab9dd15c1aff864fbb9e610066

Observation 4ee50429-81f4-44da-bb38-b3047538be34 · outbound

This paper cites Icdar2019 com- petition on scanned receipt ocr and information extraction.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Icdar2019 com- petition on scanned receipt ocr and information extraction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.533761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.533761Z digest=sha256:3769c75e3093aaa2ebad92099c371146395855e6dff814e8e5bd0ab5086ba6cf

Observation 03ad07f8-1f7b-4ddc-a91e-4d83bd6cd882 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.538690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.538690Z digest=sha256:be2350729dc31018eadcc288b7590350bdaab2f8bca3845939eb2e21897ceee3

Observation e8b21873-11d7-4150-9eab-11a29c6f2f12 · outbound

This paper cites T-rex2: Towards generic object detec- tion via text-visual prompt synergy.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding T-rex2: Towards generic object detec- tion via text-visual prompt synergy

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.544189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.544189Z digest=sha256:803e316a39b9146617d5eb960989e8c68803dda58ee848f6a124754bcaef46ea

Observation 1c5d74d5-ab14-4d63-a8a2-bdb9adb071da · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.549879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.549879Z digest=sha256:51737028d92bbf1e6f0ff978f02ea0791e1db754a17082b7e22b956729a7f695

Observation a964daef-5fd3-44d6-9fe8-982e2ecf46a5 · outbound

This paper cites A diagram is worth a dozen images.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding A diagram is worth a dozen images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.554995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.554995Z digest=sha256:a27fc5b2f4e3b967c1a6a14520da22597431f51383ddc361588f7561b091e343

Observation e0fab3ea-1bdd-4494-a132-ab39c3ed0183 · outbound

This paper cites Segment Anything.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Segment Anything

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.560592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.560592Z digest=sha256:286405c8cd1fd96de940f7b4b51995e5e16c33c6260fb8dd2a7e987ae5f3e946

Observation 7ce005b6-582b-4185-ba00-ebbf54d339f4 · outbound

This paper cites The open images dataset v4: Unified image classifica- tion, object detection, and visual relationship detection at scale.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding The open images dataset v4: Unified image classifica- tion, object detection, and visual relationship detection at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.565668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.565668Z digest=sha256:ec88cc9ac7ea67ee9ce8ae9042d0c492bd2d3e128ddd49181ff6cd6f3a0ea1b7

Observation abd9a563-b0ef-4a01-992e-f118bbf77b27 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LISA: Reasoning Segmentation via Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.570163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.570163Z digest=sha256:bc7a21d19c7b31b6434c561b1623af97a26ff75bb9dc4a70d6e62210c30365db

Observation 3a47fddb-585e-4bde-9989-2ed609e05e5f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.575794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.575794Z digest=sha256:b8ff938b34bdea8657a4b5598210f9e7481a04bf9e30cc85abfa567ce2a08e68

Observation 206cc295-fd15-4fb3-aad2-fc8a00480be8 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.581483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.581483Z digest=sha256:1ed5b7bee0b11440f20cde0e2e787d83a3459636b2a13ef21dc80cfcccce6e18

Observation 34022ed7-6c7b-4b9b-b58e-1d0a0da3565f · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.586763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.586763Z digest=sha256:e902f33c9a766a4998932ea7610210b642bdf6150fff7b20b195e18466325bd8

Observation 802e4187-9296-4023-832b-b491efc50e33 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.592398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.592398Z digest=sha256:d2775ae4ae1eb02d1f9008c8b9a4b87894976e9f303e96f19a724b4afcb30c06

Observation e63d1198-7f0c-4b6e-bc8d-e84d9693c4cc · outbound

This paper cites Grounded language-image pre-training.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounded language-image pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.597634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.597634Z digest=sha256:e0eb025063de96e9dfa2361bf7d2e965f9d5d134695bb0d4e3b432ea614faf92

Observation e376fe7e-29cf-4f28-b173-783ef072a483 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Evaluating object hallucination in large vision-language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.603708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.603708Z digest=sha256:4b6a88b0f454c41c3d82d2f60cb344f577154453b346d7cd6cb6e83e7dc9ece4

Observation 80263683-1793-451d-b747-a8208f0cc706 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.614251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.614251Z digest=sha256:692c6eb57616683fc0d03a85c344cefdb3617aa193a43fc1dd16c66828eba52e

Observation 55873e8b-ff74-4978-b429-92c8535b0f80 · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.619612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.619612Z digest=sha256:e3b66a756035dc5d28cb3554779ac1214cbede848e0097653bd8b4e574405d07

Observation 03250065-3e06-4edf-a99c-4269ecb23042 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.625353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.625353Z digest=sha256:6c0b845a44d9b9994d957d7770c3d63b33f59f2fd0fbe6ebd0a3ac7b318dd79b

Observation 245b8a8b-6018-45b7-89fb-3618ebecf326 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Vila: On pre-training for vi- sual language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.630977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.630977Z digest=sha256:c5ba589b430928335509a95bedaa28496507e1fc049db1ab22d70c98de4c3575

Observation 8ea6f3b6-3f0c-4e31-8145-d806fae5eade · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.636393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.636393Z digest=sha256:3143cf9ae8d5c5ee37f34d573f7912e275f9920381a6d77dff5fc53d39163261

Observation 53f9e16a-ded3-4fcd-ae61-c5b315355dbf · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.641744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.641744Z digest=sha256:0ce9c67eef880b7c426c7e4ef9c97c521dc4ec9dc8b14cfd3c24a4efba68fff4

Observation 376387ea-5c5f-4575-87a3-8facbff5244d · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.647571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.647571Z digest=sha256:b637d1b7ee04bb13481eaa7a1f9e8fb8af5f6cc1945d4e680da4a344bfce90a3

Observation bbf46bb7-cd75-4769-b4ad-1f09d8a9e86e · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Improved Baselines with Visual Instruction Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.652553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.652553Z digest=sha256:1e429f8710ae26a1294c3245bc3a2bb691bfd1d7c85ba30430522861b1d36d94

Observation fefe776c-5636-4920-95f0-f98fa39dcada · outbound

This paper cites Visual instruction tuning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Visual instruction tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.658293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.658293Z digest=sha256:63665cbffb0b0c1057ce8018be3cca66db6b5f3a3a32088f3089fd8fd4d31fb4

Observation e72ebb23-5906-4e13-b8c3-e6bbe368bb3b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.663513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.663513Z digest=sha256:d2d8dc197dfc736b60a547a0156b745ebd3b427b01fb6c6621590f2746bde11b

Observation 6a7e0f3e-3217-41eb-8e70-680987f1715e · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.668807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.668807Z digest=sha256:4a9c426e48d3e6d03eb5b55a6bbf7a28fff31564434f5c9016f71baeed4b1738

Observation 054af2fd-d678-4085-9848-0c951120dfd5 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.674747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.674747Z digest=sha256:af1ef86330f53d939bcb182330be983fdb1eaf5661579ca340c8ffdee844f631

Observation f0ab74b2-b8dc-4fa1-b349-6a84fd32af55 · outbound

This paper cites Ocrbench: on the hidden mys- tery of ocr in large multimodal models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Ocrbench: on the hidden mys- tery of ocr in large multimodal models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.679611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.679611Z digest=sha256:ddd8c679174a0d5e5c5973eb833681306b0b88be5e0e8cb0245a625f184d0077

Observation 5d7cd212-b136-426c-b6df-5f3e4774d97b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Swin transformer: Hierarchical vision transformer using shifted windows

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.684940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.684940Z digest=sha256:2038f627fd46aef1fde7df7aa8ed7d24cdf58b60aca7e417c4dbeb1fe8d6e97c

Observation 51de555d-bcc4-4124-87cf-cc1909032f9a · outbound

This paper cites A convnet for the 2020s.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding A convnet for the 2020s

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.690640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.690640Z digest=sha256:f069208963ac3998b7c5db2c924a4dbd934a93c71380447a3d5c7230ebd0d46f

Observation e184c4b4-aed2-4705-968d-2d608eb88878 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Towards end-to-end unified scene text detection and layout analysis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.695648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.695648Z digest=sha256:23836f93b74de9bd350f387b933c4b1e90cba631921b19d36546d576cabb2de3

Observation 90033afd-268b-4609-a24f-234dfbff9486 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.706382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.706382Z digest=sha256:2808a6f40f4cfcdafdec5981aba9e69c5d48423ca717638a9d6ced54dcaf7bc9

Observation ec05580b-55f9-47b3-8f33-4751a25e6efa · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding KOSMOS-2.5: A Multimodal Literate Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.711726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.711726Z digest=sha256:8ecb6eab6e0ade8b59fb4a2226405b59a500985395f2128a1103e474e786a947

Observation 12308234-30e1-44ed-814e-28e2f36fb9b3 · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.717150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.717150Z digest=sha256:5316d3a413b03a53a432b434a52f6295f857f3f4a38c3a1f73a19a190357e81c

Observation b4534563-8430-48b7-b1a8-c5a984c02213 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Generation and comprehension of unambiguous object descriptions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.327069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.722879Z digest=sha256:245ec67addb71c738eae808f7fac78c414d497d7115adcd40416e654f8d274f1

Observation 86f407a5-eaf7-469a-849a-5252c052f3ba · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.727517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.727517Z digest=sha256:daf037e221c579eacdd2439cba9fd1605a68cae488131efd9e3258d9c7833a0a

Observation fa6d93a9-769d-48be-b490-20f46a7cd755 · outbound

This paper cites Gpt-4v(ision) system card.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gpt-4v(ision) system card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.732753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.732753Z digest=sha256:ff12a8b9d88e6e8d485801bef8882714766917e7a30a0e8df1b3273c6f1b09dc

Observation 25cd984b-fce6-4e41-a3ee-6e93e8620074 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.737406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.737406Z digest=sha256:57c9cfdc084f032d69f7920d5254e357cd14d83fa780785a5f5541c2f63842a6

Observation dadb70b8-5652-4eaa-97d5-19a2cab5ec6b · outbound

This paper cites Perceptiongpt: Effectively fusing visual perception into llm.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Perceptiongpt: Effectively fusing visual perception into llm

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.299413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.742754Z digest=sha256:9fe1879746debd4c6db72beeccf4b402166b1ba5695899a850496025842514c4

Observation 695c8c57-982f-45f6-afe7-95b205ba2046 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Learning transferable visual models from natural language supervision

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.283423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.748671Z digest=sha256:e67f1f0121ff64800d966454084f87b0f278770b71d42327ef404ec57e0f8c4c

Observation 196b3d6d-c0c1-4c75-a44e-858f16d654ef · outbound

This paper cites Paco: Parts 11 and attributes of common objects.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Paco: Parts 11 and attributes of common objects

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.267026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.754204Z digest=sha256:c11cc6a2bf382e70eae7aa69551b1955e97e6ad1172b09486a2c47567c7e0cff

Observation f75b2626-cb3c-48ef-abd7-476173ae6985 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Glamm: Pixel grounding large multimodal model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.759129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.759129Z digest=sha256:3554277346cfb9e47c3de825c25dc36c5ecc79668cfa4b48add17648d2229f6b

Observation 3465c83b-5709-461b-82d7-7b104d721219 · outbound

This paper cites Girshick, and Jian Sun.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Girshick, and Jian Sun

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.241138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.764632Z digest=sha256:3e7022d176b46686151f31748a37c9d9f809055db5c73986ed9c8a4e8688eced

Observation e233440e-e0bd-4320-8839-c0e8c2d4740c · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.770035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.770035Z digest=sha256:2a516606b21d29a1c5070304c53a308e52181e208274b0b97e4a69ac75aee6cc

Observation e23fd7e0-2519-4635-9a3c-4bd43e84d4b5 · outbound

This paper cites Generalized in- tersection over union: A metric and a loss for bounding box regression.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Generalized in- tersection over union: A metric and a loss for bounding box regression

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.775985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.775985Z digest=sha256:dac530cd9461f5d1c84eecf01c57d9c5d88f3e83f098677b6d09211c933f1690

Observation 2df2c1ad-5e8f-4522-af0d-e33ed65c0154 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.216046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.780868Z digest=sha256:2cb1d5f4a3b3eed0ffde8aa2206431b6eae8ec8776057a0815c03864bedf174a

Observation b64f5af3-eca8-406c-b2d0-db2b54e7940a · outbound

This paper cites CrowdHuman: A Benchmark for Detecting Human in a Crowd.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CrowdHuman: A Benchmark for Detecting Human in a Crowd

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.786399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.786399Z digest=sha256:f031d6639cf68f837c5b01a81223c04c03fcfc64d0b53890f9e0336f0285f0b4

Observation 12e98316-3883-4cfc-9989-d427f362307c · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Objects365: A large-scale, high-quality dataset for object detection

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.200779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.792175Z digest=sha256:258189caf9841b2cbd2879501848351f0641ad7ed914c0c1b7396eda9d753d31

Observation df506477-f157-4d6d-9192-7b5a093d8d76 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.797984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.797984Z digest=sha256:b174733eb50da830c80728b87f346524935563be2262c889b781c3d803576176

Observation 0db5c0a3-e366-42fc-a107-a04da1d506f9 · outbound

This paper cites Towards vqa models that can read.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Towards vqa models that can read

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.184568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.803140Z digest=sha256:6a41d1ed543effc085e56876d2664418c216f02ac7f9225394711e37e7bc5e56

Observation 9e0afbd2-cbf9-4e2d-84d4-688df13d5798 · outbound

This paper cites Hashimoto.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Hashimoto

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.168651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.808437Z digest=sha256:3938c374f9f9f451adb52b732348d3e80a0ce4c1c68c933851b53966954f7162

Observation e0f781fb-9217-4ebc-a702-0b76aa3adebc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.813622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.813622Z digest=sha256:acc75b06b811d7ed8eda954599381984713485522e7fdea2210bb73133fe4776

Observation c29c3f7b-339a-4bcb-9a09-81316e169f4e · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Internlm: A multilingual language model with progressively enhanced capabilities

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.819373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.819373Z digest=sha256:256351f08b41765bd56d8008fa7f961a161cfd99afa4b51644722aaa5860a47f

Observation f51eaa3d-f744-476d-89b1-d1ffcf935919 · outbound

This paper cites Cambrian-1: A fully open, vision-centric ex- ploration of multimodal llms.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Cambrian-1: A fully open, vision-centric ex- ploration of multimodal llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.142772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.824318Z digest=sha256:04c9f5c66ba0819391e3d8b8622c25bb4975492169265a1a729a05bea33137d9

Observation beef477f-9c6c-46d8-a5a7-a21083fb81e4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.829556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.829556Z digest=sha256:da68e31c91eb862e59b6b6a544d4f0494a410d2981ec0ce45f8bcd13cab6c682

Observation b2c85656-9442-40bf-807e-5c882275d847 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.834775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.834775Z digest=sha256:c5a06c9284c61566205575c532dd567df12451ee24e135553898216af7e2abc4

Observation bf20ef6e-2ef1-4fba-bd1e-9ebc317301a6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.839351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.839351Z digest=sha256:2077d5f30e1a2f69190cca0bc7997a79a99ed818fce7640e56d57de44dc57c4e

Observation 30464c58-6124-4646-ada6-e74e40fd90e5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.844286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.844286Z digest=sha256:12131ebbab43ef3348138bfbf026fb0ebbc172ba71f022d3bbab8ef2c9e7c76a

Observation 585288b8-058f-4bc6-a95a-b5a3a7a589b7 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.850035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.850035Z digest=sha256:67825d43eb69b06ee0b39efb73f417cc327df1354aadaad4a88780400121b2d4

Observation b495736b-61db-4173-a1ef-e5670de323a0 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.855529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.855529Z digest=sha256:cef7ad139911f20b208bd0bc1b7893ab5090af1ead2292543b3bb53f5f8a3f8c

Observation 8dde6206-e697-4acf-959d-817503224947 · outbound

This paper cites Florence-2: Advancing a unified representation for a va- riety of vision tasks.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Florence-2: Advancing a unified representation for a va- riety of vision tasks

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.128195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.861408Z digest=sha256:c9e9bd907d82d10af33f4f71aad805c8ba00d593a17c353c03a5d2d94d242e90

Observation 33b7389b-5513-4a1d-9481-c34d3bcb9e98 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.866411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.866411Z digest=sha256:23ac76b6ccb2b0d8d329ed8bd2fafee09908cbd4d4aa9c7e63600db0b069c8c8

Observation b24a736c-29db-4215-baa7-c81b097b72d5 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.872055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.872055Z digest=sha256:6b0602ee4dbf3ee67e55ac481babe3d3839e71743cd525a13f8eff0525fb284b

Observation cead7cea-608a-445e-9452-62dd927f3976 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding xgen-mm (blip-3): A family of open large multimodal models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.877696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.877696Z digest=sha256:bde30e8fab21865df84324677f2254c9d4af5902b8f7d52eea11ca20297513f8

Observation 2e92401f-df96-4bea-9e3c-f3172154ebf9 · outbound

This paper cites Qwen2.5 Technical Report.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Qwen2.5 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.883200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.883200Z digest=sha256:7fb98080eca2481f64e0da352365376ab0d0f827b2074b4b15cddec6eff448cb

Observation 4b81d476-9c47-421c-a7eb-51cd12d2db36 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.888396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.888396Z digest=sha256:9c8a57576f06bb86c9fb30bf4ad0996e675e88f8f1c9d65884a52e96c7023901

Observation c2e09333-7e0d-484b-959b-c8401490223e · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.893651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.893651Z digest=sha256:e54f00e1d93899d6fd3a01d0b6ecb1732428f89f938c15b9875b0963327082d6

Observation 70984f70-42c8-4e13-93e4-30c46108de17 · outbound

This paper cites Modeling context in referring expres- sions.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Modeling context in referring expres- sions

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.112385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.899019Z digest=sha256:4940fde2b2ac2bc087c414e3b7fffd82d19ace05f38471a98e37a25f99f893a1

Observation 76703c34-0d78-4813-a605-f8e98b935373 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.904672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.904672Z digest=sha256:f4302986234af137c00656af3f9a1fadf7d4bd9cb0d92a47fde8236cce58f84d

Observation c24e2520-c16b-4e88-8344-794e34c4f095 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Osprey: Pixel understanding with visual instruction tuning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.094725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.909509Z digest=sha256:5e341af1686919124550eb59f5633f2c4247adee04bb1bdc0c2e5d701df76721

Observation 70e6fbe0-9b32-4049-9a34-519dc19d1341 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.914045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.914045Z digest=sha256:d7c48fac3f6e9389c30592eca9e55153fa58d2bb328d271fcfe8e5a5d029d897

Observation 4f46053d-b3fa-41d8-94e2-129f3ee57a77 · outbound

This paper cites From recognition to cognition: Visual commonsense rea- soning.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding From recognition to cognition: Visual commonsense rea- soning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.078043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.918748Z digest=sha256:3c26f6b5f0b7424285e70bdd552bead3b9b048c876765772842726bca9471e57

Observation ce95c1e9-0208-4384-a744-d0282c005580 · outbound

This paper cites Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.923271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.923271Z digest=sha256:c08dade00e794ce4526cea938a3a62fb97aa5aa74e2852f7dca46c4c4fc27394

Observation c28a04c0-5913-4f92-8649-fb0a03a5f1a2 · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Griffon: Spelling out all object locations at any granularity with large language models

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:19:35.061919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:19:33.928182Z digest=sha256:92daf454a1d0393b5931d9379058f35d8375c2f530e65cdccf1bf48c5c84d120

Observation 781524ef-e52a-4a72-8e29-d2400b40ec8f · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.933240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.933240Z digest=sha256:3bf0c7c0d0d12e2e89194774caf6393668442671af067b2c3d700f6669464c10

Pith citing papers

Observation 66437ede-5786-4b06-80fd-7bf12bb9742e · inbound

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding cites this paper.

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:47.703948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:50:47.703948Z digest=sha256:1a1ccc9b513712564a631c47220c8b429ca92ef0b8de93a7579aacce9523584a

Observation c9de07c2-fa78-4dac-b2cb-26ffb269408d · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.691428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.691428Z digest=sha256:041c44df4b3e2b689914111a5b53434506c13f3bd62b905e35e5908502da7f56

Observation 7e4fafc4-c40b-4f2c-abba-fae1b630d798 · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.735473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.735473Z digest=sha256:c32e6bef1e5f681b1886b193fe9e78b20c981e9961e964f3b498b3ff5353e831

Observation 21bcd62c-d96f-4fc7-a957-29b10de5adb9 · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.207146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.207146Z digest=sha256:624859e4a93ba011d4aea7af00380516b5837f8eb6b8449114afd68c9b2e1c83

Observation 0a4bef3e-e2f9-4880-8b30-7aaf303fa134 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.894977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:29ff16ae2eb9ee00c163c9062f80e15dea7c5f22a2ca96213dedb178636e44cc

Observation c0b42aca-156c-4d4e-a81d-bc1744630647 · inbound

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues cites this paper.

Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.238967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T04:38:06.673737Z digest=sha256:0404d8872ef17f790ee0fd284145d2d40042f578091abb664c2bac9f4a81330a

Observation 91885753-a2b4-495b-87a1-4de3389b2bfd · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.148592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:df33b748fee7bde0a6a3d9f7b187044c7ae57e63dd23384989f3f620120a2db4

Observation c321b45e-0f05-445c-9f31-6e0f27040570 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.232665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:3ecdb234c3359d8a7a57cd32f461bea47ca4bd5209f5d496408a7b758c12ec47

Observation 0ac54816-6a0b-4701-8b6b-f8cd2e3ab968 · inbound

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding cites this paper.

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:04:35.951457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T10:01:31.638894Z digest=sha256:7938e32fe290efa1b25c80f39a93aad6a4f3fd9d2f9f60d8629a15c7fb173ed3

Observation bd0898ff-794f-494b-91c6-da4a1dec1311 · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.361272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:2f939177c81d2a1c6f854554896fdd04f825974dd3c8a603d45548b705a3531c

Observation e0c9b431-5f79-46fc-9a18-833a08027fef · inbound

Foundation-Assisted Active Learning for Object Detection Annotation cites this paper.

Foundation-Assisted Active Learning for Object Detection Annotation ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T20:21:03.797710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:21:03.797710Z digest=sha256:7d48bda22a88d5ed9b3fc2a2e25c7a6def4aa84e224078a1fe838a85d3528ce1

Observation 66d972d2-c497-4c56-b402-548652603d6c · inbound

ReferTrack: Referring Then Tracking for Embodied Visual Tracking cites this paper.

ReferTrack: Referring Then Tracking for Embodied Visual Tracking ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:00:26.098703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:00:26.098703Z digest=sha256:f480e7733caea8af300fef24fd9cb2c9409ed23bb5328ef5dbad9600d8bcdcc5

Observation df1f35da-7181-45c1-9d20-854025286865 · inbound

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding cites this paper.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.613268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.613268Z digest=sha256:8ba8074bc80a4cdea414f34e47f3f42af8b34fb61a0cf8584995e9865a0eb6db