Pith. sign in

Paper Citation Record · LEDGER

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

As of 5 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 5 inbound Pith citation observations for arXiv:2605.27295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27295 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T18:23:30.681253Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:22.620353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T05:59:37.208407Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact22
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f41fbdb9-6d8a-42a2-8ef9-24b2bd34737f · outbound

This paper cites Learning transferable visual models from natural language supervision.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:79784cc61e14c7a851b90a23ceab1fed8214b1ec946cd399106cdc36059c57b4

Observation e7fac238-a33c-465b-b7e1-fe1da02bc226 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Scaling up visual and vision-language representation learning with noisy text supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:37913673d2f418a317bcb85549c5e512084a06a3462081304dce50574ac99645

Observation 213cdbed-22de-468e-a95d-2280f227c02e · outbound

This paper cites Siglip 2: Multilingual vision-language encoders with improved semantic understanding.Localization, and Dense Features, 6, 2025.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Siglip 2: Multilingual vision-language encoders with improved semantic understanding.Localization, and Dense Features, 6, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:07b202e9c99de8d74f28e965beabcf50b9341bd9756c4699f01f7186ba081221

Observation e13f9415-5761-4104-9d7a-8442f5c2f5f8 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.303603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:31e91f28bfef4f021225b279a53c560c0db6e918f8fea0d08243841dfb1b4323

Observation 5ffa7425-b8cb-455a-b801-b0885c98c391 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.296135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:8c189b53e2afee41608f0f93e1d893c8d9edccdbfe83f7ae22036921c406bedb

Observation 4f7dc28a-ebe6-45f4-820b-f073e503127f · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:207743b3f43c63755fb16a55d30954a781c950f9b8d951523cec71f55f574780

Observation 3a88d700-ef09-4f06-8da0-2b5d972c9736 · outbound

This paper cites arXiv preprint arXiv:2502.13595 , year=.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini arXiv preprint arXiv:2502.13595 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.289872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:d661324767688005e9f92df1b63f17c3a941e28194f0a13d1f6f304c3161d0ea

Observation 56d20cda-a9d5-4452-810a-7713d314e434 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.333638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:f57b0fd4d46d6e818165746c226fa57825f6f64c89ccfc84d0a1ec16d83ca1d3

Observation 7cd2decd-4774-4b03-8060-2878214813bf · outbound

This paper cites Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:50.301057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:2f4567044e6d1f00bede9ad16ee94163fc1897b01802800005e6b0be6e204f9c

Observation 7c777efa-e00a-4993-a633-bf788410d03c · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296, 2016.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Msr-vtt: A large video description dataset for bridging video and language.2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296, 2016

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:7b289e4be67cfd3ec1468bed77aa8d6427065a7f4d3127394e6dd475a70f8a28

Observation d742cfc9-1ded-4ebf-89d7-67f3f55df680 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini BERT: pre-training of deep bidirectional transformers for language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:c6438a308664adc45279b308fb3c77bc20fc82bef37facbecb265dfa4f501940

Observation ce743864-dc51-4c9f-b3bf-ad10024c35a8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.326043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:fbc5dd21f9ca85ebb068fe7ba4a1a13231c740507c9bc3fa67a4bbc07a581683

Observation e651c241-d03a-43de-8583-23768ff10d9a · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.353685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:043354757742740c5657a0b9c28d46b368ae062275debc39f38ffa0dc8bc698d

Observation 3ba63fe9-b520-487e-b783-ea1ad1556357 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.328654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:1f4d0d7abf048f0351428e2725e859b988d5a2a055bd26043c16997883637c78

Observation d52205b2-7714-43fc-b89d-24a46f94bb24 · outbound

This paper cites Gecko: Versatile Text Embeddings Distilled from Large Language Models.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Gecko: Versatile Text Embeddings Distilled from Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.284518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:254d690e75a0dbcd9b1f5ef0371f6e996f771837979082a0aa7bcf7214613531

Observation 2ad044ba-680c-46fe-8b6a-01ce385dac3f · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.306415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:55b61333dbcdfded6ec95ca7f7cb6ac9a1fe8beee9d81a78b005e73e27b0e31e

Observation d710f208-1463-4b5d-92ac-92e19ffe8607 · outbound

This paper cites Mteb: Massive text embedding benchmark.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Mteb: Massive text embedding benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:52deed908a25bc6201351c5886e8f24241ef7d078084393261332e4021270bb3

Observation bba5a809-f0aa-4dfb-b96d-a0f6cd373c29 · outbound

This paper cites Gemini Embedding: Generalizable Embeddings from Gemini.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Gemini Embedding: Generalizable Embeddings from Gemini

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.344681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:90a300f36665e8c43599682abd671fb2485e2eef8fa890c43f6ec6d03d0ffad0

Observation 0c024d48-6300-4260-ad3f-66f8ef5c3d8e · outbound

This paper cites arXiv preprint arXiv:2510.12709 (2025) 2.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini arXiv preprint arXiv:2510.12709 (2025) 2

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.339280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:fcc7709609d8e7cc8af89245ad665f5dd679329ae38a32cee0ad1a1ee0c2fdb2

Observation 84c08e4e-a02c-4620-b905-20489c6cdc01 · outbound

This paper cites Amazon nova multimodal embeddings: State-of-the-art embedding model for agentic rag and semantic search.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Amazon nova multimodal embeddings: State-of-the-art embedding model for agentic rag and semantic search

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:3ca84fc45a8c9bbbd3ebc7f6db72f0853d0f3312be2bceb925bc4821ed41f005

Observation 7ef45a18-3a16-4182-b019-468cb531e09f · outbound

This paper cites MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.317846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:8fad24a65aa9141f99481920437c60e7824c81ae84fb806c8b3abbc78328eeff

Observation dc643c07-c8c5-47a8-a898-d3d1f5ace4d0 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.323587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:05242e5569677b53b193f04c5188bfc417817ce152c58730c7ab4b53f7cd1889

Observation 03687859-8d4e-430f-a72f-9b69413fbeb8 · outbound

This paper cites Adapting Decoder-Based Language Models for Diverse Encoder Downstream Tasks.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Adapting Decoder-Based Language Models for Diverse Encoder Downstream Tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.350783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:def90ea66ba9908a2dc3effa625c67b8da126039f9aa781cddf504670cfe962f

Observation 8422d2df-a23c-4a7a-af62-5ea36896a9f4 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.298511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:4a38e5ae1febcbdf4ac202c0c0ac1ad573d976ab59fd8b59c5fa63fe40ba9836

Observation 334b8734-373d-41bd-bd77-5439f0bd8a97 · outbound

This paper cites Matryoshka representation learning.Advances in Neural Information Processing Systems, 35:30233–30249, 2022.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Matryoshka representation learning.Advances in Neural Information Processing Systems, 35:30233–30249, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:f4a486fc32ad59455fa5af46806b6a1a83afd996f1565ab4cf4f79a6891cbcf6

Observation db4929d4-fdfb-40b6-aae9-6d4454025633 · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Averaging Weights Leads to Wider Optima and Better Generalization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.287012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:f1bc0797f81b521066b15dc6d75e18978c034247592aa2242c502810fc7169c3

Observation 34daf114-b575-4c37-a88d-5776a2211aab · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:4be71d814b37a41ea9923c51f494baa32646fcb4af067b356851be7ea4060489

Observation 30204292-9702-4283-8913-7ec4349a7b9d · outbound

This paper cites Introducing the google uni- versal image embedding challenge.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Introducing the google uni- versal image embedding challenge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:f8994093d9ca5066872bf88f26a5e5c5b56b3032cec83182a62f6c4e752e4240

Observation 127b89fe-f73c-4f12-99bb-77eb594c1e52 · outbound

This paper cites author Dong, W.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini author Dong, W

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.165452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:7e4adbd200fe163cb18f4f6d1d907fd240971e5fa432eff041f7382c8d855181

Observation 3269f5d8-5b06-4512-80fc-667a87290235 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini DOCCI: Descriptions of Connected and Contrasting Images

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.309525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:24c9d28fb622513ca5d1abfd2d99c41e31fe6d1ca209143af8d81238ae5fa5ad

Observation 1bce227a-59d6-415e-b499-4256eaefe95c · outbound

This paper cites TextCaps: A Dataset for Image Captioning with Reading Comprehension.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini TextCaps: A Dataset for Image Captioning with Reading Comprehension

Reference 31

Resolution
verified exact
doi, observed 2026-06-29T18:23:50.160109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:2045e0fec18c12151ca6b6bf7a717344709a209305139e7df4f4512624878921

Observation 38df4655-2de9-4185-9099-1b5ea4b23d27 · outbound

This paper cites VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.342053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:211c22a84353045a6ee6826e44205233db1907c63d20bf931252407e8f453aab

Observation f534cd35-885a-48bd-949a-f64023e4143b · outbound

This paper cites Towards Automatic Learning of Procedures from Web Instructional Videos.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Towards Automatic Learning of Procedures from Web Instructional Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.311975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:92961ca76e4593b837b5b299709eabd26a497bc03aef6ced9292786c417476df

Observation bc15209d-c146-4a70-8b1d-5e164b8b9d4e · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:92e1f9246ef7cf2687e712ecee1b14ddd885e4f02fab62ecc24da4df14e4dea5

Observation 2a0846f5-b618-4c4c-b5b9-924da8970d6a · outbound

This paper cites Vidore benchmark v2: Raising the bar for visual retrieval.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Vidore benchmark v2: Raising the bar for visual retrieval

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.320704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:556947b427a32ebb0decbf2748c197fff610738a3388d2df6ce51fdeb1a071f6

Observation c0b38e93-8fe6-477b-9f06-777a70241b14 · outbound

This paper cites Voyage multimodal 3.5.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Voyage multimodal 3.5

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:e340dc995554703dffefd600495e584865fe49a33a1bab8d17c3f3b730bc3bbc

Observation 40102aaa-6fa9-4226-99ac-3cf50918c7d1 · outbound

This paper cites Multimodal embeddings API.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Multimodal embeddings API

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:ec8271d448e90c9f777c9981d0aaed306928e803371e87e03f227e416f16d848

Observation 5298879f-654e-4063-8c00-c3f3a4b1dfeb · outbound

This paper cites Google universal image embedding.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Google universal image embedding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:3e46f925aeac28390ed3f3b61c02af61e7e490e07938ace152f2512eb0335d14

Observation b7fefc0c-c373-46c0-ae36-a56182e55700 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Msr-vtt: A large video description dataset for bridging video and language

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:782d3c2f4d18d3bcfcc1112c49f3ee4f1f9434d4698868861ee9da96a62747b8

Observation d99d0576-0d7a-4a9f-8708-0057fb070238 · outbound

This paper cites CoIR: A Comprehensive Benchmark for Code Information Retrieval Models.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini CoIR: A Comprehensive Benchmark for Code Information Retrieval Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.292747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:5292c382da6dc811ed273e7bf50de4726fd73e2445bcfb5a41dc8c2443df49c3

Observation 62109dfa-d410-4f52-b3e8-71c70c08b56f · outbound

This paper cites Massive sound embedding benchmark (mseb), 2026.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Massive sound embedding benchmark (mseb), 2026

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.331337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:0a0576a8a45d50c47df38bdc5799a79015f7bc57070494340669dd591cf519c2

Observation 47673b0a-9f6b-4a96-a8a9-c3290244c129 · outbound

This paper cites Microvqa: A multimodal reasoning benchmark for microscopy-based scientific research.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Microvqa: A multimodal reasoning benchmark for microscopy-based scientific research

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:160ebe04e6f041cb906c9f5e29a3bf5d69e218ae9403f8201a4fc0f370312a28

Observation d96bfdf5-0b7f-4955-b849-53ae9f7e01dd · outbound

This paper cites Artcap: A dataset for image captioning of fine art paintings.IEEE Transactions on Computational Social Systems, 11(1):576–587, 2022.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Artcap: A dataset for image captioning of fine art paintings.IEEE Transactions on Computational Social Systems, 11(1):576–587, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:07619fe8b2deef8a36b0357f91b464d73755b7a354a2c36e89e90e0317f4aa90

Observation 15f4af27-8fe3-42f4-96f2-0a4e65a08b7f · outbound

This paper cites AstroLLaVA: towards the unification of astronomical data and natural language.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini AstroLLaVA: towards the unification of astronomical data and natural language

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.347887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:b34de72aec6c4dea94448ec887f9175f23e30fb5f28687a93f96903796ee5352

Observation 915b2211-651a-4a3e-a19e-67a44cf31621 · outbound

This paper cites Recipe1m+: A dataset for learning cross-modal embeddings for cooking recipes and food images.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1):187–203, 2021.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Recipe1m+: A dataset for learning cross-modal embeddings for cooking recipes and food images.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1):187–203, 2021

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:3f4838dad81bd4901048d37be341796a4b5fc5d5baad1c304407cf86d2d9d855

Observation ee8dc2f1-8720-4862-87a8-d8a4171b509c · outbound

This paper cites TIPS: Text-Image Pretraining with Spatial awareness.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini TIPS: Text-Image Pretraining with Spatial awareness

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.314604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:45a4f63f99b156f5bacbf6e02df3df2fc789f7481a65497e3b97edf38bedb638

Observation aa90389d-2a0e-4e6e-a67e-dcf99a679f1d · outbound

This paper cites Opencodeinterpreter: Integrating code generation with execution and refinement,.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Opencodeinterpreter: Integrating code generation with execution and refinement,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:7b15adb023829f426b0baf3d876c1abb595a7150206adb596c93c8d93251ea52

Observation a3b673c5-9590-426d-aa02-a43b1018e3b7 · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.336541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:2e399c1fd60f24171517e687bea852b2cd8c919fbc54721fe2fa183a6226c3c8

Observation 4eecb4c8-6f00-46e7-99ed-566dcd2f5c26 · outbound

This paper cites an unresolved cited work.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:fed91dca6eeaf4582dcb2a0aa735e8aca587dc0e402e24d528b336b4b2b20b2f

Observation b184677e-fcf6-494f-875f-104ef8ef4b6c · outbound

This paper cites an unresolved cited work.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T18:23:30.681253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:625e2ea068cd35dc1bb62c1526fe48480e3e88491ac26a87c2b74b5d1b55e4d6

Pith citing papers

Observation 00cb9db0-bb9c-4779-b138-f5072ca87c96 · inbound

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment cites this paper.

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:07:37.108120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T14:08:00.387995Z digest=sha256:a97f2ac8373a3c76335352dab94aed386ee295b574b1d496392935a273110195

Observation 00a20ab4-c19d-4a05-b381-d5bb9d3652b5 · inbound

The Token Tax of Epistemic Accuracy: Comparing RAG and Long-Context Architectures for Document-Grounded Generative AI Applications cites this paper.

The Token Tax of Epistemic Accuracy: Comparing RAG and Long-Context Architectures for Document-Grounded Generative AI Applications Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:59:37.210157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T15:08:16.105336Z digest=sha256:591de1e252ce602af5ff0f8c57052e30ab57ed7073e41c2c6df6d231b47f492d

Observation ac8ec57b-f516-4392-a998-a3e457e54e34 · inbound

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes cites this paper.

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T19:22:55.637577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:22:55.637577Z digest=sha256:fb36a4dbaf97fd8655efaf427130e27925a871afeb463984e4924706ed605105

Observation 404724dd-83f2-4b20-ab67-34c4e15f0bc6 · inbound

Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating cites this paper.

Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T03:59:33.933416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:59:33.933416Z digest=sha256:b5912186c75be1434315430f62b895794cf46daf359d9a85b5daa4f4a1608c1c

Observation e83e423f-e064-4aec-96ef-b12ca0841479 · inbound

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding cites this paper.

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:22.620353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:43:22.620353Z digest=sha256:4738dafe20da22c255897078a1b7c038f44dd1f7db1254f4d3590bd08bd25e81