Pith. sign in

Paper Citation Record · LEDGER

Region-based Cluster Discrimination for Visual Representation Learning

As of 9 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2507.20025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20025 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:56:27.154410Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:45:23.828033Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:25:11.499059Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy58
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d27e9e5-90cd-47e1-b314-216d3e546ae3 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Region-based Cluster Discrimination for Visual Representation Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.620832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.620832Z digest=sha256:ec8b9084771c9fbae75615564d62a9fd00957952f25ec2be47dd078aec2f5b93

Observation 8868673a-26eb-4cf2-b49b-0c61982daff4 · outbound

This paper cites Killing Two Birds with One Stone: Efficient and Robust Training of Face Recogni- tion CNNs by Partial FC.

Region-based Cluster Discrimination for Visual Representation Learning Killing Two Birds with One Stone: Efficient and Robust Training of Face Recogni- tion CNNs by Partial FC

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.626364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.626364Z digest=sha256:27e7b0aa06c9aa5df4dc0c4694f29fb6230d667bbde846cd0a5ebb9a4f2bc9f4

Observation 23c24aa5-7378-4b86-90bb-b65e2b4d4185 · outbound

This paper cites Unicom: Universal and Compact Representation Learning for Image Retrieval.

Region-based Cluster Discrimination for Visual Representation Learning Unicom: Universal and Compact Representation Learning for Image Retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.631180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.631180Z digest=sha256:7d444a9bf1fdf945ccb60faad038594ff25227f32232c5ae2acd814d18abc2dd

Observation 3df184d6-0d62-46b5-9918-15d2ce409e0c · outbound

This paper cites Multi-label Cluster Discrimination for Vi- sual Representation Learning.

Region-based Cluster Discrimination for Visual Representation Learning Multi-label Cluster Discrimination for Vi- sual Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.637509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.637509Z digest=sha256:01430b721bf579db53c824e963687e6164a9f984eeec5e4208115fe1e4cd6314

Observation c4bf6569-b086-45ac-999b-3e4c1faa0cbe · outbound

This paper cites Self-Labelling via Simultaneous Clustering and Representation Learning.

Region-based Cluster Discrimination for Visual Representation Learning Self-Labelling via Simultaneous Clustering and Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.642982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.642982Z digest=sha256:2c3d93a4aaac8f7e821be135d15d7e90742d633021caa993f9be084a5ad1d65f

Observation a2dfd4bc-a44f-409b-9d59-2b3fc8f9a51c · outbound

This paper cites Qwen technical report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen technical report

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.066111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.651574Z digest=sha256:676d3934b7c12e47c50b5f6f8b4862e4cc9926c4b1b46657403d800a460c2c56

Observation e2820eea-58b2-499b-a168-dadc71350e12 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Region-based Cluster Discrimination for Visual Representation Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.656817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.656817Z digest=sha256:3ed774efd37b3914a18f72b6be6a766c01d604c54cf47077b8dc8869c6d6af4b

Observation 6b2cdd11-877a-4c4d-8371-ce279011b5aa · outbound

This paper cites COYO-700M: Image-Text Pair Dataset, 2022.

Region-based Cluster Discrimination for Visual Representation Learning COYO-700M: Image-Text Pair Dataset, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.046242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.662383Z digest=sha256:f8df66c2083dcc4607c1429ec214823a71eec0f0546d6f49d9ffdf090a45e938

Observation f108fb0a-08be-4dfb-a16d-d819fab94525 · outbound

This paper cites Cascade R-CNN: Delv- ing Into High Quality Object Detection.

Region-based Cluster Discrimination for Visual Representation Learning Cascade R-CNN: Delv- ing Into High Quality Object Detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.026478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.666694Z digest=sha256:ca7539c93c41079190e5b96f1f04192ed508a148287373fe249692981fb46b7f

Observation 0fffa4a4-41a0-4c5e-a078-b6f749525d48 · outbound

This paper cites Deep Clustering for Unsupervised Learn- ing of Visual Features.

Region-based Cluster Discrimination for Visual Representation Learning Deep Clustering for Unsupervised Learn- ing of Visual Features

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.998554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.671354Z digest=sha256:9f940269d850602b47361df14d1f25efe6d06436b974eb49ea396ccaeb804852

Observation 578b62d4-a23a-46e4-a847-3b1d28b6a4cd · outbound

This paper cites Unsupervised Learning of Visual Features by Contrasting Cluster Assignments.

Region-based Cluster Discrimination for Visual Representation Learning Unsupervised Learning of Visual Features by Contrasting Cluster Assignments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.978389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.676512Z digest=sha256:11a639d34883008f8425a2bf094815f7b9ae1a45ca621c56a81a9310c90972a5

Observation 8e9e1edb-639c-4061-bfeb-20baea1782d5 · outbound

This paper cites ViTamin: Designing Scalable Vision Models in the Vision-language Era.

Region-based Cluster Discrimination for Visual Representation Learning ViTamin: Designing Scalable Vision Models in the Vision-language Era

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.955934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.681641Z digest=sha256:3ca75a4388eab759ef922ecfaac95b876a4d65368fe20b353f7378a14532459b

Observation 1e2a0df4-f371-4e92-b130-f2c8ec6c88ff · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? In NeurIPS,.

Region-based Cluster Discrimination for Visual Representation Learning Are We on the Right Way for Evaluating Large Vision-Language Models? In NeurIPS,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.929358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.686659Z digest=sha256:3c70614bf61e08d1eee4499f5c25f298f5e78249267b41c9d04313ec67eb9550

Observation c3b4a854-0718-4eeb-a6ed-0b934ad31461 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Region-based Cluster Discrimination for Visual Representation Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.692234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.692234Z digest=sha256:b0bf2d4591a968f9ca280e4a730f9227a76a2e067eea5090931d951a4fdd4ce5

Observation bfa25025-d2fe-41ed-a6e3-979cb33f21a4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Region-based Cluster Discrimination for Visual Representation Learning Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.907699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.697377Z digest=sha256:e8bb0ddd71a380b9659d6baa2ccdd427b529762567d12310d19be506a0340a66

Observation 7dfce565-8f2b-4ad1-93c3-541aa343f1a4 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Region-based Cluster Discrimination for Visual Representation Learning Arcface: Additive angular margin loss for deep face recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.886230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.702288Z digest=sha256:ae091d7956ade074e011e7de9a9a287da9dbf2cd227824f3a0dcfede6d9b1143

Observation d8324a6a-4407-4564-8ed6-ddb1548f8dd2 · outbound

This paper cites The Faiss library.

Region-based Cluster Discrimination for Visual Representation Learning The Faiss library

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.706863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.706863Z digest=sha256:13b4927147ea7b192c05ebc65ca6594ae721a746a5a0d41b9d1a5afc820f0aa9

Observation b146aa89-7293-4f63-b417-42e088c20978 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

Region-based Cluster Discrimination for Visual Representation Learning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.712695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.712695Z digest=sha256:59d508fc65b1ce63b7b0ece447460f1530ea0619837524f3dc86f2e8f9987bec

Observation 29539a18-5b0b-4e5d-8870-d4fba7cc9c45 · outbound

This paper cites Susskind, and Armand Joulin.

Region-based Cluster Discrimination for Visual Representation Learning Susskind, and Armand Joulin

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.867030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.717823Z digest=sha256:97313a5e99de59c305033aa0b53ffeb5341bf984b55c0545adec9915c56125fb

Observation c9476024-a4a2-4b42-8843-323ab6014549 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Region-based Cluster Discrimination for Visual Representation Learning Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.846191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.724305Z digest=sha256:768d2720d53956979202ded39f2ee3763e447608e69e3d296e1a52917f87e7bb

Observation dd485a37-c933-4ca5-b1fa-67ff4b57e0d5 · outbound

This paper cites Data Filtering Networks.

Region-based Cluster Discrimination for Visual Representation Learning Data Filtering Networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.823850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.730472Z digest=sha256:378cbc785ebc24690f5544abba6971e4e6d6c281d1738da8978ec8331ece8b77

Observation 7da053c8-a2d1-410d-b8fa-7375d0bfc7ff · outbound

This paper cites Multimodal autoregres- sive pre-training of large vision encoders.

Region-based Cluster Discrimination for Visual Representation Learning Multimodal autoregres- sive pre-training of large vision encoders

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.798405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.735675Z digest=sha256:1bd93bcdfec09c0d7b0ee679314f45b1bcdcb3f3f86396cad220b812c015a0fc

Observation 9eca4821-f4c1-4417-92bb-4244ff285e9c · outbound

This paper cites Rwkv-clip: A robust vision-language representation learner.

Region-based Cluster Discrimination for Visual Representation Learning Rwkv-clip: A robust vision-language representation learner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.779670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.740772Z digest=sha256:7135fbf417cec763a9ad5fa1fd26153045be81cfce9feb9fdcb7de5304560432

Observation a33c675a-2da5-4bb6-9a0c-2f8318c0da50 · outbound

This paper cites Mask R-CNN.

Region-based Cluster Discrimination for Visual Representation Learning Mask R-CNN

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.753046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.746235Z digest=sha256:7cd7a960bcbc1ea427e1d5f025252b033de96c7f581cfb1939dd4287fbc292a1

Observation 53731b7c-aa3f-4b22-ab30-b6e47e40e7f4 · outbound

This paper cites GOT-10k: A Large High-Diversity Benchmark for Generic Object Track- ing in the Wild.

Region-based Cluster Discrimination for Visual Representation Learning GOT-10k: A Large High-Diversity Benchmark for Generic Object Track- ing in the Wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.730690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.750708Z digest=sha256:ef693d1403b41489890862dcd074277e62985c0daa68a89b674d6439b3caaca7

Observation 831c0921-b0f0-4ddf-8bbb-63248ddfcd90 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2.5-Coder Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.755498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.755498Z digest=sha256:466e0f7d8b75097108fb1835dec1443976969e52c504b7b3f24c817278684e89

Observation 1f99e61b-8935-4477-8fae-7a8ce65b8268 · outbound

This paper cites Open- CLIP.

Region-based Cluster Discrimination for Visual Representation Learning Open- CLIP

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.700746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.760557Z digest=sha256:6f4e5207e4df9497fbabd0848c8b2c16cbce04abbfdecd568d5e0fb35f5a66d6

Observation 25b17057-5684-4168-a833-b837c1a7f00b · outbound

This paper cites ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes.

Region-based Cluster Discrimination for Visual Representation Learning ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.674280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.765785Z digest=sha256:c8d8674d1d29ff55e0bacda4febabb6242eaf10a7b79d9c16d49dc7b72a75877

Observation dcd687b2-8d25-4930-b35b-8edb59ccba1f · outbound

This paper cites A Diagram Is Worth A Dozen Images.

Region-based Cluster Discrimination for Visual Representation Learning A Diagram Is Worth A Dozen Images

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.650994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.771976Z digest=sha256:eb93ff939b513d9d6eec4f40de231af9d1f11c75372d797dd505ab2e31bfc709

Observation a7fb8d0d-3d8a-4779-ac90-ffc4779320cb · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Region-based Cluster Discrimination for Visual Representation Learning Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.632191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.777420Z digest=sha256:c9705e1c49909a2299fe099a10fc9f09eb65dc48632d8952d0e5c90c50c2d62b

Observation 3e702191-796e-49c1-a3cd-030fc1a6177b · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

Region-based Cluster Discrimination for Visual Representation Learning Lisa: Reasoning segmenta- tion via large language model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.607227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.782673Z digest=sha256:478e9a8f130bb0a9fbe933915783d1f25dcf3efb625b900782760eb8fc6d6eae

Observation 7bb3a0c4-e890-42c1-a10f-d8b71f41b409 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Region-based Cluster Discrimination for Visual Representation Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.787539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.787539Z digest=sha256:3f0a6e6ca32d472361cfbcf82c27e6ee94d6c3335f90604b7c8c2e9efed5f759

Observation 9f0f6b97-000a-48ed-bc77-a3243f92d05a · outbound

This paper cites LLaV A-NeXT- Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-NeXT- Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.585248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.793074Z digest=sha256:dd28334f47ebbfb0e68609f63f1bd0408415cae8f517c5e591dbe750d4ecc741

Observation b3d192ce-99cc-4172-ab06-8e832e4deb7e · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation.

Region-based Cluster Discrimination for Visual Representation Learning BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.566597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.797553Z digest=sha256:d6ffdd1cd3bb0b90d9fdeb2236cc21f1cf5fe4ea2dc0e4f7353ca81b77b87cc8

Observation f6a983ab-5036-4cde-8167-d20eadee74bb · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.802594Z digest=sha256:d24f16df9d2bff61f603bf06b2345585025d79b3fa8bacaf6287c85dc29bbf47

Observation 1b9bec1e-7071-4909-86ed-b072b299f8f4 · outbound

This paper cites Grounded Language-Image Pre-training.

Region-based Cluster Discrimination for Visual Representation Learning Grounded Language-Image Pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.535545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.809119Z digest=sha256:a2ef5a80b5d729ae104574128582eb76f70ff7160769946c5e97636684d1886b

Observation f7bf54cd-4a72-425a-afff-4da048306f93 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Region-based Cluster Discrimination for Visual Representation Learning Evaluating Object Hallucination in Large Vision-Language Models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.518580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.815122Z digest=sha256:843cc1a821d358415c51a77e8a5d269eea21bee4f696277d1c4c69139915f5dd

Observation 7c4e8872-88e8-4dda-8e84-eff047bc3706 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Region-based Cluster Discrimination for Visual Representation Learning Improved Baselines with Visual Instruction Tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.494563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.820691Z digest=sha256:4fba3241cbdeec07ef4493feb429db153a61ef5efac05c2cfb88f8852b047d06

Observation bdf6e1e0-5c17-4ddc-8560-e01fe56d7b3a · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.474788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.827107Z digest=sha256:6ea5e866dcf07494a7595c61f523616355a52844de5e0cb101300b46badf18b5

Observation 48d63edc-1a04-4082-8c6c-3645c3df4881 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Region-based Cluster Discrimination for Visual Representation Learning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.832006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.832006Z digest=sha256:4a748c50ae6074b11a58ac24eb3e0594d5e9fb3d0d9f2979e8bd942a8fb43298

Observation 753a9abb-feb5-4ca5-a431-1860ae9fa98b · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.452881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.837827Z digest=sha256:e082610183d107e8562250a501ea6af7976dc0fcfcd26c8e0df6c1ec3cd4a1c6

Observation ce541bc0-2f36-4504-8aaf-a20c5f68913c · outbound

This paper cites A ConvNet for the 2020s.

Region-based Cluster Discrimination for Visual Representation Learning A ConvNet for the 2020s

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.427578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.845274Z digest=sha256:552177aac4fc3104bd88ca71da350c11f6df1c3cf8c2d46eaa3f3334e4361f0a

Observation 27363820-ae25-4de8-be49-c9d041d58616 · outbound

This paper cites Decoupled Weight Decay Regularization.

Region-based Cluster Discrimination for Visual Representation Learning Decoupled Weight Decay Regularization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.405520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.850988Z digest=sha256:61b5f865398af69f0c6c88da73e5b65497ebae3207a5c58a8e08a2c722769d96

Observation 8325b1ff-7101-45a9-b261-cec3714f2db6 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Region-based Cluster Discrimination for Visual Representation Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.385809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.863231Z digest=sha256:682fa2bc0967dd15daa2910eb23b9c5878299e482ba7a1e765b4b8acd7d30f7e

Observation 8a7b096b-fd37-40ff-bb55-8683b2c7e2cb · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Region-based Cluster Discrimination for Visual Representation Learning DocVQA: A Dataset for VQA on Document Images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.361845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.867868Z digest=sha256:4adc183f6cd3c2f5216f92c1c38fffd44fe0aa42fc155d565425d3c440788c81

Observation d8b543ac-7c86-416a-b3ff-6403778287c1 · outbound

This paper cites Infographicvqa.

Region-based Cluster Discrimination for Visual Representation Learning Infographicvqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.326016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.872346Z digest=sha256:6959ebde6ad55aa0e5f1e5d17ea07399a8fcc2ca771720b2d4449ac69cce3f36

Observation 7ad7ac98-048f-4fff-80fd-2ce7c300efac · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

Region-based Cluster Discrimination for Visual Representation Learning Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.305168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.876897Z digest=sha256:60781b448252aca7ef30394378e55ed730bb95543776359d0763686b85bbb285

Observation 3ccbe1b7-1617-4de5-81a7-5c8db38b6b02 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Region-based Cluster Discrimination for Visual Representation Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.282928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.886959Z digest=sha256:81a6fac187cd2b9d6d01995d1bd53833a92bed7dfebc1262b462dcd2f7d8108e

Observation 8409c45a-edc2-488d-a0c7-584227948456 · outbound

This paper cites Filtering, Dis- tillation, and Hard Negatives for Vision-Language Pre- Training.

Region-based Cluster Discrimination for Visual Representation Learning Filtering, Dis- tillation, and Hard Negatives for Vision-Language Pre- Training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.264964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.892625Z digest=sha256:269fb94164a5d935a836905d0b8ed588c32c03f735a870b4e0c5d1d9e7f97336

Observation 3a50cf7c-26f0-4d4e-97cb-92cb6e9c42d0 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Region-based Cluster Discrimination for Visual Representation Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.898477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.898477Z digest=sha256:53558ac150c38de0ac604be4af500dd225145a03883ca1b8820fa04508fc4944

Observation b5e727e8-f463-45c6-9112-b84d5741ffab · outbound

This paper cites Pre- Det: Large-Scale Weakly Supervised Pre-Training for De- tection.

Region-based Cluster Discrimination for Visual Representation Learning Pre- Det: Large-Scale Weakly Supervised Pre-Training for De- tection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.230114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.903501Z digest=sha256:da2b79ea4a84fb357f8d945b38717e0cb2046ca23de747e139b036fdd78ac8b2

Observation c10f915b-06bd-417d-b452-4dfa7ea248b9 · outbound

This paper cites GLaMM: Pixel Grounding Large Multimodal Model.

Region-based Cluster Discrimination for Visual Representation Learning GLaMM: Pixel Grounding Large Multimodal Model

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.212070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.908430Z digest=sha256:95596d16144aac376fd442955bcc321d8ee2c903e905a7a17743a11bd28b3157

Observation e52daf83-4f0e-4f25-baa8-f8689b1cebd1 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Region-based Cluster Discrimination for Visual Representation Learning You Only Look Once: Unified, Real-Time Object Detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.192059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.917252Z digest=sha256:5ea8c89d3837bf846486ee78567c1df3e51cbb4f1bc985ca3787cedb28a14ddc

Observation be276d48-b716-4cf1-93ad-d34f601e39ad · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Region-based Cluster Discrimination for Visual Representation Learning Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.173309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.923646Z digest=sha256:b6c6a71fa80b20d41b88294c48102c4fc04fd98992c23bed4ef3d38ff1eec09c

Observation 530b4c3e-d5af-4c39-b87b-1646dfe36edc · outbound

This paper cites PixelLM: Pixel Reasoning with Large Multimodal Model.

Region-based Cluster Discrimination for Visual Representation Learning PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.155359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.928907Z digest=sha256:a98e1b2f0d8affeab43df4d8e48b60a6332c50e60bd2fee3c4bac5b17f7cdcea

Observation 2b5aff1b-a1e1-4b92-894c-dcd9865ccf87 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Region-based Cluster Discrimination for Visual Representation Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.933437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.933437Z digest=sha256:1d2795d4657375abf03a66f485907d0e85bca349f69cabdd65388f6cb30d8ec6

Observation b0ed33ef-de76-4f7c-befc-f9cf0b90222e · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Region-based Cluster Discrimination for Visual Representation Learning LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.939298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.939298Z digest=sha256:9d138abec563288250eb5c3c64db0f4626427a15cd125f0a7b5e196d31cf3236

Observation b60b6483-f0a3-4a0d-a0c9-a35b9243415f · outbound

This paper cites LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content.

Region-based Cluster Discrimination for Visual Representation Learning LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:56:27.520942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.945884Z digest=sha256:7fde1a910cdb90aea06230e3a266311e0f770c79af2bd76f0fa569c18071818b

Observation 7a4a02a5-98a9-4564-9c87-bf97628ac1d1 · outbound

This paper cites Towards VQA Models That Can Read.

Region-based Cluster Discrimination for Visual Representation Learning Towards VQA Models That Can Read

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.134453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.951466Z digest=sha256:e51db54bce44d14d0c982ea2638c729d68d8ff8c9df940f156e9e7954f8e9106

Observation 5ba26a7c-3263-4440-bbb7-fba95592b1b8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Region-based Cluster Discrimination for Visual Representation Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.956775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.956775Z digest=sha256:09319258c0294e4174fad056b20a08f28801872f04243a63b5332b1ba8f41b97

Observation e8f3085e-e9a6-4d88-bfa3-2b10bd6d99c8 · outbound

This paper cites InternLM2 Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning InternLM2 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.970571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.970571Z digest=sha256:38a61fd98372ee3899cc91949a19670b8688c1731ed155b928d4357654908525

Observation dc6c820e-b7ed-4718-ad31-292029e152f4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Region-based Cluster Discrimination for Visual Representation Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.976368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.976368Z digest=sha256:54b9ed1f2d3046c0d8154640dcd163d4e9d69291ba67a6e682bb389185a04fa6

Observation d35d2446-da5c-4306-98b9-b9a8b73a0c16 · outbound

This paper cites Qwen2 Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.980964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.980964Z digest=sha256:7c85769cfcad5157395a2c8b70b3dbe791fdbd9aff30151015b03016312bbecb

Observation cfebeb82-92bf-4a0c-8c1b-d47e28915a94 · outbound

This paper cites Qwen2.5-VL Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.985779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.985779Z digest=sha256:c7888c4c4a7a4ea25374e02f4950a9b762059ada488657544cb5721d91595991

Observation bf795244-f07b-47d3-a853-564efd7d6380 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Ex- ploration of Multimodal LLMs.

Region-based Cluster Discrimination for Visual Representation Learning Cambrian-1: A Fully Open, Vision-Centric Ex- ploration of Multimodal LLMs

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.111533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.991478Z digest=sha256:9647681bbdfe5aba112debccac2431db8f68d0c95f803adb53a6f1064ff1045c

Observation 3819060f-3dc4-4dfc-8c04-30c8e84759d0 · outbound

This paper cites SigLIP 2: Multilingual Vision- Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Region-based Cluster Discrimination for Visual Representation Learning SigLIP 2: Multilingual Vision- Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.091143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.001638Z digest=sha256:4a0b66cccfdd56546f565d5e793ff8a386d98bc6b8e97650c3450008a5e239c6

Observation 6d6b98e1-9908-4a3e-ba3a-6f43545751f2 · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

Region-based Cluster Discrimination for Visual Representation Learning Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.071897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.011523Z digest=sha256:a4658902ba8742ffd39824af1b7285478013ea42b57a0fa2958620e8bc084b65

Observation 657d9363-1684-4ead-a37b-93345585e330 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Region-based Cluster Discrimination for Visual Representation Learning LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.018502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.018502Z digest=sha256:66d0df5fc523fb288dc0d7dc210f3cb3976e617f93a43b7c7b4696760063b087

Observation f4823ff2-3b92-4f85-b405-7b6a4de8ad42 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Region-based Cluster Discrimination for Visual Representation Learning VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.053255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.028344Z digest=sha256:7e3449e805623e28d5f1cfd5709cf2d1c6c73a95b2bdde9e6a05883be3af9430

Observation acda5a5a-9103-49b0-b146-687cec0e5dec · outbound

This paper cites CLIM: Contrastive Language-Image Mosaic for Region Representation.

Region-based Cluster Discrimination for Visual Representation Learning CLIM: Contrastive Language-Image Mosaic for Region Representation

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:56:27.300988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.035845Z digest=sha256:56937a85ead0c4389a7a817b98b71bb175a65b066bab69532f36fd53c27f1850

Observation b1af74ed-cb42-455f-acb4-b746dbf93ecb · outbound

This paper cites Detectron2, 2019.

Region-based Cluster Discrimination for Visual Representation Learning Detectron2, 2019

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.024377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.043462Z digest=sha256:4b99ab82af925e8524a8163a908174a9b39cf2f6fac854ec7b86cd67d899f597

Observation 201015b5-7acc-4979-b2f7-a1bcb35e7faa · outbound

This paper cites Grok-1.5 Vision Preview, 2024.

Region-based Cluster Discrimination for Visual Representation Learning Grok-1.5 Vision Preview, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.002476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.049705Z digest=sha256:419cc26de6d52b4b54eaa0ec4a509a23848e5ed432933c58ab87a72809058083

Observation f61d3b58-7c8c-4d2a-9fcc-9ece30bbb3cb · outbound

This paper cites GSV A: Generalized Segmentation via Multimodal Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning GSV A: Generalized Segmentation via Multimodal Large Language Models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.981192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.054660Z digest=sha256:45a2318b8eb9d80f582e3a55ee7a83b472933f39112842ff3e52c4f2a5cb316d

Observation f83317b3-ff39-41fd-be81-9d7f03d44781 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic cap- tion.

Region-based Cluster Discrimination for Visual Representation Learning Alip: Adaptive language-image pre-training with synthetic cap- tion

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.957878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.060124Z digest=sha256:bceb1081dd2ef3eb158193d041c4edff416883b8ea1f5e011a37b1e28e291575

Observation 9c9f9909-f995-4054-84fa-73af2da37528 · outbound

This paper cites Clip-cid: Efficient clip distillation via cluster-instance discrimination.

Region-based Cluster Discrimination for Visual Representation Learning Clip-cid: Efficient clip distillation via cluster-instance discrimination

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.936687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.066531Z digest=sha256:7dd850284c61262ec339be9624866c7ff1b43960de703b107749f9b7bd88c790

Observation 9cc1d279-fad1-46d3-96c0-a32e559c724d · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

Region-based Cluster Discrimination for Visual Representation Learning Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.915607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.073703Z digest=sha256:23de39f3c8a24976296cbb37800c04b2b237b29944b9865707b30c83411a102d

Observation 339569e8-c32b-4f03-b790-3ac816ade7f1 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning A Survey on Multimodal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.081336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.081336Z digest=sha256:c9bd5bbff9049db180055f6bef9f64beecc10a673db296002f0ebdb18430b025

Observation 66bad569-f09b-4ece-bc1c-a3277640fc09 · outbound

This paper cites Berg, and Tamara L.

Region-based Cluster Discrimination for Visual Representation Learning Berg, and Tamara L

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.893374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.088735Z digest=sha256:df6edde7a116679901a0cb7e63055b6319c86f3efe1c4453763e10a1dea61807

Observation 3bc442f7-61bb-497c-94ee-a3eacdfbc400 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Region-based Cluster Discrimination for Visual Representation Learning Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.094504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.094504Z digest=sha256:dfa60645b5ea1e17a332b8c9f5ed77b9a736e556708594b868c93127f949f202

Observation e2b9d908-9dc4-4cce-b933-a0198cb1c23d · outbound

This paper cites Sigmoid Loss for Language Image Pre- Training.

Region-based Cluster Discrimination for Visual Representation Learning Sigmoid Loss for Language Image Pre- Training

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.872321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.104031Z digest=sha256:d8a105c5149b4e20bdabca3d6e9e2a2e8327d0c5ba4dd4c3b64e10dacd87a1e7

Observation e49f113f-e79b-416e-98d8-2837d5d0bfe2 · outbound

This paper cites LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.847480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.122836Z digest=sha256:d190bcb5376edd9a7fef065954f855092018f2ccee37f48fe1fa310df53a647b

Observation 5673a67a-59b7-45a3-97f2-e0392912f68e · outbound

This paper cites OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding.

Region-based Cluster Discrimination for Visual Representation Learning OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.818735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.130640Z digest=sha256:c3ec42d85dc8db9a66cd6c74ef1de17f51132c38e8d58e7f14f22ab48863baf9

Observation 2120dd39-0480-4a4d-a14f-4c6458c60d2f · outbound

This paper cites RegionCLIP: Region-based Language-Image Pretraining.

Region-based Cluster Discrimination for Visual Representation Learning RegionCLIP: Region-based Language-Image Pretraining

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.794744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.141777Z digest=sha256:92ab541b284e88be3a1f703c0a42e5ab119eb4f9d4ca6d03d1d843a7171a6d8d

Observation 3253df0b-b858-4373-9757-f289e7f7a7e5 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

Region-based Cluster Discrimination for Visual Representation Learning MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.148197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.148197Z digest=sha256:c8678d9eb17330bc6c6ad3f511609becb7587bf2be6e4d16b38a0db3dfa32e53

Observation 2c432c4c-a465-4244-aaec-79ecb7ecaa98 · outbound

This paper cites UNIT: Unifying Image and Text Recognition in One Vision Encoder.

Region-based Cluster Discrimination for Visual Representation Learning UNIT: Unifying Image and Text Recognition in One Vision Encoder

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.761473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.154410Z digest=sha256:f05380d72e920b2760f7afd31fa0999c99c0e688dd4f16c3643b79e4e1b16369

Pith citing papers

Observation fda919ed-db74-41ad-bab1-53962652ed2a · inbound

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval cites this paper.

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval Region-based Cluster Discrimination for Visual Representation Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:23.828033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:23.828033Z digest=sha256:0580d139900b6c7eefc63485b912d99128d0778b50796ef69830638609369c5d

Observation 94f0dd6d-908b-4968-9f9a-ccadb3d367a6 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts Region-based Cluster Discrimination for Visual Representation Learning

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.501302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:5efb8478faa75842f58cf28c02d25cd41f5e973142e96ba317386429a857f6ee