Pith. sign in

Paper Citation Record · LEDGER

Region-based Cluster Discrimination for Visual Representation Learning

As of 9 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2507.20025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20025 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:56:27.154410Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:45:23.828033Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:25:11.499059Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy58
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d27e9e5-90cd-47e1-b314-216d3e546ae3 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Region-based Cluster Discrimination for Visual Representation Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.620832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.620832Z digest=sha256:5a0cc8903d19350601b0dababcaed52f71136339eff342ea89d6bac9c18fdaa1

Observation 8868673a-26eb-4cf2-b49b-0c61982daff4 · outbound

This paper cites Killing Two Birds with One Stone: Efficient and Robust Training of Face Recogni- tion CNNs by Partial FC.

Region-based Cluster Discrimination for Visual Representation Learning Killing Two Birds with One Stone: Efficient and Robust Training of Face Recogni- tion CNNs by Partial FC

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.626364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.626364Z digest=sha256:80b376f52fc9d47d9edd4456e90073f86dd68e102678488e9bebcd1eedfc664f

Observation 23c24aa5-7378-4b86-90bb-b65e2b4d4185 · outbound

This paper cites Unicom: Universal and Compact Representation Learning for Image Retrieval.

Region-based Cluster Discrimination for Visual Representation Learning Unicom: Universal and Compact Representation Learning for Image Retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.631180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.631180Z digest=sha256:38848743358cdc546dfead722488c43a3545f7644166a15e0a1f067f1eb0fa0a

Observation 3df184d6-0d62-46b5-9918-15d2ce409e0c · outbound

This paper cites Multi-label Cluster Discrimination for Vi- sual Representation Learning.

Region-based Cluster Discrimination for Visual Representation Learning Multi-label Cluster Discrimination for Vi- sual Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.637509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.637509Z digest=sha256:fea4fee9695d631a918e3dcb5ff6e9b24afbb5a008c0592b915ab3fb6982870b

Observation c4bf6569-b086-45ac-999b-3e4c1faa0cbe · outbound

This paper cites Self-Labelling via Simultaneous Clustering and Representation Learning.

Region-based Cluster Discrimination for Visual Representation Learning Self-Labelling via Simultaneous Clustering and Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.642982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.642982Z digest=sha256:ec3156d39f226e067dbfec9e4300e6e2ba19831fb31a6c31689fb018a858999f

Observation a2dfd4bc-a44f-409b-9d59-2b3fc8f9a51c · outbound

This paper cites Qwen technical report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen technical report

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.066111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.651574Z digest=sha256:4551788f4402ffdc4832fa8b62f9359a3101b0e9e9ac113b5de2979f228bc549

Observation e2820eea-58b2-499b-a168-dadc71350e12 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Region-based Cluster Discrimination for Visual Representation Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.656817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.656817Z digest=sha256:f75dc91d1ee76db9e0c58afbeea7e7359dcccb5b93e829e644512f5cfb32e9d1

Observation 6b2cdd11-877a-4c4d-8371-ce279011b5aa · outbound

This paper cites COYO-700M: Image-Text Pair Dataset, 2022.

Region-based Cluster Discrimination for Visual Representation Learning COYO-700M: Image-Text Pair Dataset, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.046242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.662383Z digest=sha256:ae6e14adf8dbc1ac51d8f73c0bd6995557610f7e211787d03b5b8f48f5f179a1

Observation f108fb0a-08be-4dfb-a16d-d819fab94525 · outbound

This paper cites Cascade R-CNN: Delv- ing Into High Quality Object Detection.

Region-based Cluster Discrimination for Visual Representation Learning Cascade R-CNN: Delv- ing Into High Quality Object Detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:29.026478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.666694Z digest=sha256:c2922718a5727e3c161486b13cf7827db7c59b48d6fc8a84189985891ce9b8b2

Observation 0fffa4a4-41a0-4c5e-a078-b6f749525d48 · outbound

This paper cites Deep Clustering for Unsupervised Learn- ing of Visual Features.

Region-based Cluster Discrimination for Visual Representation Learning Deep Clustering for Unsupervised Learn- ing of Visual Features

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.998554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.671354Z digest=sha256:a590abd54473c251891d1cc3570979c033e4960c77c104e27412bac279e62e0a

Observation 578b62d4-a23a-46e4-a847-3b1d28b6a4cd · outbound

This paper cites Unsupervised Learning of Visual Features by Contrasting Cluster Assignments.

Region-based Cluster Discrimination for Visual Representation Learning Unsupervised Learning of Visual Features by Contrasting Cluster Assignments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.978389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.676512Z digest=sha256:466133670c7f0f0df233c96e9686fe5849287f17d9a78358156e8c70bcac2894

Observation 8e9e1edb-639c-4061-bfeb-20baea1782d5 · outbound

This paper cites ViTamin: Designing Scalable Vision Models in the Vision-language Era.

Region-based Cluster Discrimination for Visual Representation Learning ViTamin: Designing Scalable Vision Models in the Vision-language Era

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.955934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.681641Z digest=sha256:9fa3dc5718f3d2b702fb8c373f700df1072d279e0bdbc65f6fb9eccc9b3b779d

Observation 1e2a0df4-f371-4e92-b130-f2c8ec6c88ff · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? In NeurIPS,.

Region-based Cluster Discrimination for Visual Representation Learning Are We on the Right Way for Evaluating Large Vision-Language Models? In NeurIPS,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.929358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.686659Z digest=sha256:73c4479bb8d0c8bac139248d80dce11fb732488825d8d33f640838e7d7a6c1d3

Observation c3b4a854-0718-4eeb-a6ed-0b934ad31461 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Region-based Cluster Discrimination for Visual Representation Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.692234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.692234Z digest=sha256:41dc9b723d63c9748edd525e31e8039ba824617dbce74f8054e438632c890d3f

Observation bfa25025-d2fe-41ed-a6e3-979cb33f21a4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Region-based Cluster Discrimination for Visual Representation Learning Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.907699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.697377Z digest=sha256:dcba3c39141196f933ed8eb2699b72212d62b90866b99dbe83e1e2e06b8455ea

Observation 7dfce565-8f2b-4ad1-93c3-541aa343f1a4 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Region-based Cluster Discrimination for Visual Representation Learning Arcface: Additive angular margin loss for deep face recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.886230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.702288Z digest=sha256:8510dad406d6620dd77a762c356faf0f70e5f5a29ccd20796b75cd978ee7f5d0

Observation d8324a6a-4407-4564-8ed6-ddb1548f8dd2 · outbound

This paper cites The Faiss library.

Region-based Cluster Discrimination for Visual Representation Learning The Faiss library

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.706863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.706863Z digest=sha256:a639e3c2a9c8926f2bcab47ad0ce19af702476ac87df52e6e8f0b2c1d8493928

Observation b146aa89-7293-4f63-b417-42e088c20978 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

Region-based Cluster Discrimination for Visual Representation Learning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.712695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.712695Z digest=sha256:36c2c6f26d629a8cc5d752119d25484849a37bbf4bc18bfd418994d694a72c9f

Observation 29539a18-5b0b-4e5d-8870-d4fba7cc9c45 · outbound

This paper cites Susskind, and Armand Joulin.

Region-based Cluster Discrimination for Visual Representation Learning Susskind, and Armand Joulin

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.867030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.717823Z digest=sha256:84962cdafc7ee81a5260f1ade81610d58f963850737eb5db779d0936364458a3

Observation c9476024-a4a2-4b42-8843-323ab6014549 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Region-based Cluster Discrimination for Visual Representation Learning Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.846191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.724305Z digest=sha256:359aa326391419fba31148e91ead247b697df338543bd600510d23508362e0e6

Observation dd485a37-c933-4ca5-b1fa-67ff4b57e0d5 · outbound

This paper cites Data Filtering Networks.

Region-based Cluster Discrimination for Visual Representation Learning Data Filtering Networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.823850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.730472Z digest=sha256:b24c079734235ae386399f5adc4f476e57958fe8db6bbfedcd8c8653164f732e

Observation 7da053c8-a2d1-410d-b8fa-7375d0bfc7ff · outbound

This paper cites Multimodal autoregres- sive pre-training of large vision encoders.

Region-based Cluster Discrimination for Visual Representation Learning Multimodal autoregres- sive pre-training of large vision encoders

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.798405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.735675Z digest=sha256:3f151abd8734caf0f0df023bca6f4afd7679cd687cd2662b24f30153f7072f8b

Observation 9eca4821-f4c1-4417-92bb-4244ff285e9c · outbound

This paper cites Rwkv-clip: A robust vision-language representation learner.

Region-based Cluster Discrimination for Visual Representation Learning Rwkv-clip: A robust vision-language representation learner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.779670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.740772Z digest=sha256:505a3fa7ffde7575040eaa1d271ffc7d5767cd94ad58ae73190079073206f24b

Observation a33c675a-2da5-4bb6-9a0c-2f8318c0da50 · outbound

This paper cites Mask R-CNN.

Region-based Cluster Discrimination for Visual Representation Learning Mask R-CNN

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.753046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.746235Z digest=sha256:32fe8c5a4c7795728f60750965951539f42855d33accfc0c508a59c52d426610

Observation 53731b7c-aa3f-4b22-ab30-b6e47e40e7f4 · outbound

This paper cites GOT-10k: A Large High-Diversity Benchmark for Generic Object Track- ing in the Wild.

Region-based Cluster Discrimination for Visual Representation Learning GOT-10k: A Large High-Diversity Benchmark for Generic Object Track- ing in the Wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.730690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.750708Z digest=sha256:0c58261072a72e281411f3dcd5c8f08ca3eafb2c67cc46fd8ae4058f5e4a63b3

Observation 831c0921-b0f0-4ddf-8bbb-63248ddfcd90 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2.5-Coder Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.755498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.755498Z digest=sha256:3078c34fc42103335eec80e922bd66bfe31f97c568ee7e81ff8482cd326cbebc

Observation 1f99e61b-8935-4477-8fae-7a8ce65b8268 · outbound

This paper cites Open- CLIP.

Region-based Cluster Discrimination for Visual Representation Learning Open- CLIP

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.700746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.760557Z digest=sha256:323d3dc0565e67b8821d7b12ce991ef4bb31f083d9c55235ddf0ee2f4f293a31

Observation 25b17057-5684-4168-a833-b837c1a7f00b · outbound

This paper cites ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes.

Region-based Cluster Discrimination for Visual Representation Learning ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.674280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.765785Z digest=sha256:ecd541498fe0e7252bb3e6cb97f44b23807aeeb95400e1bc067e31bb04bff162

Observation dcd687b2-8d25-4930-b35b-8edb59ccba1f · outbound

This paper cites A Diagram Is Worth A Dozen Images.

Region-based Cluster Discrimination for Visual Representation Learning A Diagram Is Worth A Dozen Images

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.650994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.771976Z digest=sha256:8183a461e5ea65d99a717db791226be1acedc68eb4a986592ebc7182836db686

Observation a7fb8d0d-3d8a-4779-ac90-ffc4779320cb · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Region-based Cluster Discrimination for Visual Representation Learning Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.632191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.777420Z digest=sha256:ccb228236a95cabdc03c76220278309a3ed426b0b1d61adfb8f76fad72da142b

Observation 3e702191-796e-49c1-a3cd-030fc1a6177b · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

Region-based Cluster Discrimination for Visual Representation Learning Lisa: Reasoning segmenta- tion via large language model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.607227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.782673Z digest=sha256:5c084e22e613c1c61db84698369162967366de8be0686e5448dedf94b92b7509

Observation 7bb3a0c4-e890-42c1-a10f-d8b71f41b409 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Region-based Cluster Discrimination for Visual Representation Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.787539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.787539Z digest=sha256:be404a7c432e1b3bef57c347c9bfc8baaa316c5265a0fd3a141bd2edf09106f2

Observation 9f0f6b97-000a-48ed-bc77-a3243f92d05a · outbound

This paper cites LLaV A-NeXT- Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-NeXT- Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.585248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.793074Z digest=sha256:e8d1859acb945ab3eb7dc2c17d7fc08980f0993cdf38d7d0898232b1a6818785

Observation b3d192ce-99cc-4172-ab06-8e832e4deb7e · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation.

Region-based Cluster Discrimination for Visual Representation Learning BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.566597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.797553Z digest=sha256:389a921fbda0b380350adb793f6a0bc7b9b8129358696ad2967d52ee7f380bd0

Observation f6a983ab-5036-4cde-8167-d20eadee74bb · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.802594Z digest=sha256:09724a0af395cb11c41ced86e4f11149e9440ead652d4e55191bafc3e8ab433e

Observation 1b9bec1e-7071-4909-86ed-b072b299f8f4 · outbound

This paper cites Grounded Language-Image Pre-training.

Region-based Cluster Discrimination for Visual Representation Learning Grounded Language-Image Pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.535545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.809119Z digest=sha256:c72733d77500e9791735229d19ded9e335bf83b14a07cab845c51d266cc438db

Observation f7bf54cd-4a72-425a-afff-4da048306f93 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Region-based Cluster Discrimination for Visual Representation Learning Evaluating Object Hallucination in Large Vision-Language Models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.518580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.815122Z digest=sha256:8591c143537774fd5e9d4fe3507a6a3582b658d8e09ac272268c1dda315b9f05

Observation 7c4e8872-88e8-4dda-8e84-eff047bc3706 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Region-based Cluster Discrimination for Visual Representation Learning Improved Baselines with Visual Instruction Tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.494563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.820691Z digest=sha256:bb72c68806b062b263e69cc0e8ab1c251f107b61cf1d2d0a5e8d22b2e158c0cf

Observation bdf6e1e0-5c17-4ddc-8560-e01fe56d7b3a · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.474788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.827107Z digest=sha256:aa089a2b4aef9e6f9f6b5d71d9c8b1819073067e49e7da5d7ca6461107d402eb

Observation 48d63edc-1a04-4082-8c6c-3645c3df4881 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Region-based Cluster Discrimination for Visual Representation Learning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.832006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.832006Z digest=sha256:30144a02c2d1a8a834763cbf2df62f7be3e4a437fd25a3c7431d99a96dd91fc1

Observation 753a9abb-feb5-4ca5-a431-1860ae9fa98b · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.452881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.837827Z digest=sha256:0e17ca06eac9fe6fc4dc5999f534363944a8ca6760e151f0ad103cf74038c821

Observation ce541bc0-2f36-4504-8aaf-a20c5f68913c · outbound

This paper cites A ConvNet for the 2020s.

Region-based Cluster Discrimination for Visual Representation Learning A ConvNet for the 2020s

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.427578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.845274Z digest=sha256:1f1958ade1a14ae07493c15fcfe1a9276a6511b4e29475346630d5197a4ede7b

Observation 27363820-ae25-4de8-be49-c9d041d58616 · outbound

This paper cites Decoupled Weight Decay Regularization.

Region-based Cluster Discrimination for Visual Representation Learning Decoupled Weight Decay Regularization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.405520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.850988Z digest=sha256:5d0e07709051951548abcbb01a66fdeb93a0d20abfa9b373ff89a743800375f2

Observation 8325b1ff-7101-45a9-b261-cec3714f2db6 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Region-based Cluster Discrimination for Visual Representation Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.385809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.863231Z digest=sha256:4aa27b691c64cec270f73f3907cb63b5d70c9f790d0eb0e734448efe0ba73276

Observation 8a7b096b-fd37-40ff-bb55-8683b2c7e2cb · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Region-based Cluster Discrimination for Visual Representation Learning DocVQA: A Dataset for VQA on Document Images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.361845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.867868Z digest=sha256:2288e553800e746aa655046b4a720aed8d35043591aeb9398b21bb8f77833b7a

Observation d8b543ac-7c86-416a-b3ff-6403778287c1 · outbound

This paper cites Infographicvqa.

Region-based Cluster Discrimination for Visual Representation Learning Infographicvqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.326016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.872346Z digest=sha256:d8049e6ede173157294ecb7230f1f49beaa8e68948036267d5092c1a3c0c7be0

Observation 7ad7ac98-048f-4fff-80fd-2ce7c300efac · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

Region-based Cluster Discrimination for Visual Representation Learning Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.305168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.876897Z digest=sha256:690108a0235ab3ef636dd857e36d076d03b0ff6e8a91b0b962522f113b390301

Observation 3ccbe1b7-1617-4de5-81a7-5c8db38b6b02 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Region-based Cluster Discrimination for Visual Representation Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.282928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.886959Z digest=sha256:c87812a80fd382d43e7d94c93a91e48f32a0c366dc2e666cf5131c2ab0fca244

Observation 8409c45a-edc2-488d-a0c7-584227948456 · outbound

This paper cites Filtering, Dis- tillation, and Hard Negatives for Vision-Language Pre- Training.

Region-based Cluster Discrimination for Visual Representation Learning Filtering, Dis- tillation, and Hard Negatives for Vision-Language Pre- Training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.264964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.892625Z digest=sha256:18d439e4d691ed7eba5a7e761fb26da57b53c3e9f08c917b71296b99e18d4b8c

Observation 3a50cf7c-26f0-4d4e-97cb-92cb6e9c42d0 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Region-based Cluster Discrimination for Visual Representation Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.898477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.898477Z digest=sha256:f6e0e8f6242e23236e8a45c15113d865f489e465416201e71f2ab510671c4390

Observation b5e727e8-f463-45c6-9112-b84d5741ffab · outbound

This paper cites Pre- Det: Large-Scale Weakly Supervised Pre-Training for De- tection.

Region-based Cluster Discrimination for Visual Representation Learning Pre- Det: Large-Scale Weakly Supervised Pre-Training for De- tection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.230114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.903501Z digest=sha256:18067e70548a43ae3cafa2dd16963a716a9b10c1b38e82f861aa3f00208709f1

Observation c10f915b-06bd-417d-b452-4dfa7ea248b9 · outbound

This paper cites GLaMM: Pixel Grounding Large Multimodal Model.

Region-based Cluster Discrimination for Visual Representation Learning GLaMM: Pixel Grounding Large Multimodal Model

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.212070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.908430Z digest=sha256:b4f9d514e5fcad8c58168814c7351ea9d157339d3ec7ae5d37974117ea5b9a83

Observation e52daf83-4f0e-4f25-baa8-f8689b1cebd1 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Region-based Cluster Discrimination for Visual Representation Learning You Only Look Once: Unified, Real-Time Object Detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.192059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.917252Z digest=sha256:86d3912a3bb4011f5d0148bc2dce8891ed7a23725362f9ca829b3836d407ee36

Observation be276d48-b716-4cf1-93ad-d34f601e39ad · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Region-based Cluster Discrimination for Visual Representation Learning Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.173309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.923646Z digest=sha256:d0abd25b39b50b76f3619f7b162f9c1696b32add6403d4425bec279d6d54d368

Observation 530b4c3e-d5af-4c39-b87b-1646dfe36edc · outbound

This paper cites PixelLM: Pixel Reasoning with Large Multimodal Model.

Region-based Cluster Discrimination for Visual Representation Learning PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.155359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.928907Z digest=sha256:72766d9844204202a8d65580822bbde004b762d066137aa1c963003a3e93c52a

Observation 2b5aff1b-a1e1-4b92-894c-dcd9865ccf87 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Region-based Cluster Discrimination for Visual Representation Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.933437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.933437Z digest=sha256:022975b02b6183adeff5801f177800620349c4e12343a4c20d33f63f7ca26d33

Observation b0ed33ef-de76-4f7c-befc-f9cf0b90222e · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Region-based Cluster Discrimination for Visual Representation Learning LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.939298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.939298Z digest=sha256:dab85171949961d1669f0130d3343aec0a55725d86ed00651c47ed7696378690

Observation b60b6483-f0a3-4a0d-a0c9-a35b9243415f · outbound

This paper cites LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content.

Region-based Cluster Discrimination for Visual Representation Learning LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:56:27.520942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.945884Z digest=sha256:36b9aaa5522c11b63c68cf5ef87f3b7e8f1e2e8cbbfa98365bdd5f477b966a9f

Observation 7a4a02a5-98a9-4564-9c87-bf97628ac1d1 · outbound

This paper cites Towards VQA Models That Can Read.

Region-based Cluster Discrimination for Visual Representation Learning Towards VQA Models That Can Read

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.134453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.951466Z digest=sha256:8a984c1a6558d89788f7b95c5df9ab2cbc1b76c03a313080d07f6a3e6344aadb

Observation 5ba26a7c-3263-4440-bbb7-fba95592b1b8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Region-based Cluster Discrimination for Visual Representation Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.956775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.956775Z digest=sha256:073da74c8bf04d3f776186ab3c6b9ed74ed45f9d8ef72e9245f96def153fd99d

Observation e8f3085e-e9a6-4d88-bfa3-2b10bd6d99c8 · outbound

This paper cites InternLM2 Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning InternLM2 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.970571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.970571Z digest=sha256:53bcccba43642a90b51507a6f644dd760ef05c6c7d709628f0661f2051894937

Observation dc6c820e-b7ed-4718-ad31-292029e152f4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Region-based Cluster Discrimination for Visual Representation Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.976368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.976368Z digest=sha256:5b5cd707f91008aef42fc348964a8b613365c52740fa861fa8fff38d46f89f11

Observation d35d2446-da5c-4306-98b9-b9a8b73a0c16 · outbound

This paper cites Qwen2 Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.980964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.980964Z digest=sha256:5fb85839bb8be9bd093f4303fc1df48c2795df1826d8cabcb0cc0b2c54804c0a

Observation cfebeb82-92bf-4a0c-8c1b-d47e28915a94 · outbound

This paper cites Qwen2.5-VL Technical Report.

Region-based Cluster Discrimination for Visual Representation Learning Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:26.985779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:26.985779Z digest=sha256:eb2f9099195f66001ef39c9dff5247db61fb859cef009a6ea7831eeca7498162

Observation bf795244-f07b-47d3-a853-564efd7d6380 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Ex- ploration of Multimodal LLMs.

Region-based Cluster Discrimination for Visual Representation Learning Cambrian-1: A Fully Open, Vision-Centric Ex- ploration of Multimodal LLMs

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.111533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:26.991478Z digest=sha256:b6a77b59b4d8a5a52c668a7d3020669273f001241a1971c34415f48460ea12d2

Observation 3819060f-3dc4-4dfc-8c04-30c8e84759d0 · outbound

This paper cites SigLIP 2: Multilingual Vision- Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Region-based Cluster Discrimination for Visual Representation Learning SigLIP 2: Multilingual Vision- Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.091143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.001638Z digest=sha256:44866e3cab9f389e99d5ea338170d2c3313a77a6d50f8ecf7b5aae2e178ac9bf

Observation 6d6b98e1-9908-4a3e-ba3a-6f43545751f2 · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

Region-based Cluster Discrimination for Visual Representation Learning Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.071897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.011523Z digest=sha256:202e0544a7b8bd55e2a669ffeb52e881816b7a96c5c27a56bfead80f9979addf

Observation 657d9363-1684-4ead-a37b-93345585e330 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Region-based Cluster Discrimination for Visual Representation Learning LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.018502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.018502Z digest=sha256:b2e48e978aff9430d674917a1e4f04d47f7b5db124ffeef012da63a86d6d8032

Observation f4823ff2-3b92-4f85-b405-7b6a4de8ad42 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Region-based Cluster Discrimination for Visual Representation Learning VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.053255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.028344Z digest=sha256:92d71c12c1a43b80afa065d55aeb817312b7b3d90939332d5a89a593ce55891b

Observation acda5a5a-9103-49b0-b146-687cec0e5dec · outbound

This paper cites CLIM: Contrastive Language-Image Mosaic for Region Representation.

Region-based Cluster Discrimination for Visual Representation Learning CLIM: Contrastive Language-Image Mosaic for Region Representation

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:56:27.300988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.035845Z digest=sha256:9d6b0486e19636bf9cbd432907ff7593a1149d3b7e6c27226c272e4b62065fe2

Observation b1af74ed-cb42-455f-acb4-b746dbf93ecb · outbound

This paper cites Detectron2, 2019.

Region-based Cluster Discrimination for Visual Representation Learning Detectron2, 2019

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.024377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.043462Z digest=sha256:5b5a4214566142de7744193fa341f0ba50fd09830ef391741b92d7822459149a

Observation 201015b5-7acc-4979-b2f7-a1bcb35e7faa · outbound

This paper cites Grok-1.5 Vision Preview, 2024.

Region-based Cluster Discrimination for Visual Representation Learning Grok-1.5 Vision Preview, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:28.002476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.049705Z digest=sha256:171e05063670f244f8d895c83ea6c4657f03ef8e32a3b8ac21fe446e0c54fddc

Observation f61d3b58-7c8c-4d2a-9fcc-9ece30bbb3cb · outbound

This paper cites GSV A: Generalized Segmentation via Multimodal Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning GSV A: Generalized Segmentation via Multimodal Large Language Models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.981192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.054660Z digest=sha256:caaef896e261d601c110c12a05ecebb8203badd488f7863b454f4f9c2aaa7e9e

Observation f83317b3-ff39-41fd-be81-9d7f03d44781 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic cap- tion.

Region-based Cluster Discrimination for Visual Representation Learning Alip: Adaptive language-image pre-training with synthetic cap- tion

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.957878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.060124Z digest=sha256:0c5c2d00e97494af375174ee6c8dc9279ed3727feae70e82704c6ef61092c50f

Observation 9c9f9909-f995-4054-84fa-73af2da37528 · outbound

This paper cites Clip-cid: Efficient clip distillation via cluster-instance discrimination.

Region-based Cluster Discrimination for Visual Representation Learning Clip-cid: Efficient clip distillation via cluster-instance discrimination

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.936687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.066531Z digest=sha256:382dab8ebba485c83a398be66c8107da76ee89eab2d4db7c2f7c50492e87164b

Observation 9cc1d279-fad1-46d3-96c0-a32e559c724d · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

Region-based Cluster Discrimination for Visual Representation Learning Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.915607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.073703Z digest=sha256:836795de6892f6084abc935e7d4aea97bd61608f318e9ffb529443205a77cb5a

Observation 339569e8-c32b-4f03-b790-3ac816ade7f1 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Region-based Cluster Discrimination for Visual Representation Learning A Survey on Multimodal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.081336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.081336Z digest=sha256:29711c8da8f67f694f865a34b4fac5756dd9698cd2d4eb646e6d346523f89ce7

Observation 66bad569-f09b-4ece-bc1c-a3277640fc09 · outbound

This paper cites Berg, and Tamara L.

Region-based Cluster Discrimination for Visual Representation Learning Berg, and Tamara L

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.893374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.088735Z digest=sha256:0d80d2ccb800f67e58702d5a9e2e05132ae8fb2243dfa58f284ad9f3af49a46c

Observation 3bc442f7-61bb-497c-94ee-a3eacdfbc400 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Region-based Cluster Discrimination for Visual Representation Learning Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.094504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.094504Z digest=sha256:114b9d4491b581294b0f3ae96fb89eda4d7c549fb80838e0d4e5283b7b9fb6e3

Observation e2b9d908-9dc4-4cce-b933-a0198cb1c23d · outbound

This paper cites Sigmoid Loss for Language Image Pre- Training.

Region-based Cluster Discrimination for Visual Representation Learning Sigmoid Loss for Language Image Pre- Training

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.872321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.104031Z digest=sha256:04de0982c82648c8341d7e0af0cd168705f2e51fc99e2b237cfee2a2c03f3616

Observation e49f113f-e79b-416e-98d8-2837d5d0bfe2 · outbound

This paper cites LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models.

Region-based Cluster Discrimination for Visual Representation Learning LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.847480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.122836Z digest=sha256:6bcc9f8128053d6126fb5ed7e8f4c55eb5bed9be5dbe318b4be7f11136b0a271

Observation 5673a67a-59b7-45a3-97f2-e0392912f68e · outbound

This paper cites OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding.

Region-based Cluster Discrimination for Visual Representation Learning OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.818735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.130640Z digest=sha256:79e845c911474866b027a8eb27287ce4a8ee5b15cb3974a80506341de5af43e0

Observation 2120dd39-0480-4a4d-a14f-4c6458c60d2f · outbound

This paper cites RegionCLIP: Region-based Language-Image Pretraining.

Region-based Cluster Discrimination for Visual Representation Learning RegionCLIP: Region-based Language-Image Pretraining

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.794744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.141777Z digest=sha256:7c59b2256e0f76be0e07ff58d3614236d2efeecd5ac5d31e5b808352a10f9776

Observation 3253df0b-b858-4373-9757-f289e7f7a7e5 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

Region-based Cluster Discrimination for Visual Representation Learning MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:27.148197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:27.148197Z digest=sha256:5fb4eadb7bdb1b9c6f99b8dac43abe451b638c7aa894b2226ac417355abf11e8

Observation 2c432c4c-a465-4244-aaec-79ecb7ecaa98 · outbound

This paper cites UNIT: Unifying Image and Text Recognition in One Vision Encoder.

Region-based Cluster Discrimination for Visual Representation Learning UNIT: Unifying Image and Text Recognition in One Vision Encoder

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:56:27.761473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:56:27.154410Z digest=sha256:ebc065abb0e495648779eb6a9931bc3cd8ea1a7fe4b97e4325ee19b81c360a02

Pith citing papers

Observation fda919ed-db74-41ad-bab1-53962652ed2a · inbound

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval cites this paper.

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval Region-based Cluster Discrimination for Visual Representation Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:23.828033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:23.828033Z digest=sha256:6bc4d3da530a122f722d14c2f22daf00997dcd13285322bbd621b76e4bb886b7

Observation 94f0dd6d-908b-4968-9f9a-ccadb3d367a6 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts Region-based Cluster Discrimination for Visual Representation Learning

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.501302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:f98ed36b4554de321c15e5cddcc475deb9f0199d00fc37b734308bd40e248ceb