Pith. sign in

Paper Citation Record · LEDGER

UNITER: UNiversal Image-TExt Representation Learning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:1909.11740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.11740 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:12:08.524191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:44:57.758784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 629aeb46-409d-49bc-9c0a-18157c1721af · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training UNITER: UNiversal Image-TExt Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.242260Z digest=sha256:5acf607ea26ec6f36d9015aff6fe1d29592a04d603b072cc1201f92d212cb1c5

Observation 5468228b-b5b5-4972-891c-b5da5a2b711a · inbound

Language Models are Few-Shot Learners cites this paper.

Language Models are Few-Shot Learners UNITER: UNiversal Image-TExt Representation Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:38.144764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:05:38.045330Z digest=sha256:431192cd3559b3358659dfcee1b7b239ac6565d265d4d2f6a958178cf4da459b

Observation 0903af5b-c090-436b-9f7f-fb006c8b7a09 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space UNITER: UNiversal Image-TExt Representation Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.424338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:6438a878a0f87d12fe0e9c0ee3bf0eb94d08f5438ba2b5c8f15b63fe3a235ca7

Observation be956813-70c8-4907-803b-68d82d645b5f · inbound

Everything is a Video: Unifying Modalities through Next-Frame Prediction cites this paper.

Everything is a Video: Unifying Modalities through Next-Frame Prediction UNITER: UNiversal Image-TExt Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.046717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.046717Z digest=sha256:859cc5165bce53f46e5d735ad4d22b614eac20410bece11ea7ab0a91c5fde322

Observation 0d2363b8-29c3-4e49-8ad5-570f7f323a1a · inbound

Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment cites this paper.

Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment UNITER: UNiversal Image-TExt Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:34:47.038879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:34:47.038879Z digest=sha256:75970eb423e92d2cc1523784f46b8d0b205ec227158ae684ced50e21a8347346

Observation 19852c09-dec0-4569-89f8-f635ef29bc6e · inbound

Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection cites this paper.

Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection UNITER: UNiversal Image-TExt Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:56:38.212564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:56:38.212564Z digest=sha256:13f5475c74facd91ef030e2ecf676038aeb40ee7621b9a1bebf7c8619fcb6731

Observation 671a0c85-5eee-4c68-bcdb-5b3249cb232b · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UNITER: UNiversal Image-TExt Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.486497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.486497Z digest=sha256:82c60457834a97369416e2ffb8313076a6a4e6f4a7e94167554ed94b6b5ae4aa

Observation 71f8190f-6cb2-4dec-818f-fcf540a2c64d · inbound

LA-RCS: LLM-Agent-Based Robot Control System cites this paper.

LA-RCS: LLM-Agent-Based Robot Control System UNITER: UNiversal Image-TExt Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:05.354331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:05.354331Z digest=sha256:ee8b9fb78154df62d43915a814b84f68bb56b2efa90ccae160dde09c02f8bff9

Observation 97cb7be6-a06c-463a-8359-e2eb0820f02c · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation UNITER: UNiversal Image-TExt Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.281943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.281943Z digest=sha256:133ea38f1992d754f3fbd521d76659b6976fc7cd6142919ddf9363533d7c85c9

Observation fde2c972-902f-483c-8a35-38d25d02358d · inbound

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation cites this paper.

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation UNITER: UNiversal Image-TExt Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:20.365794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:20.365794Z digest=sha256:90363c79eee748da634c707f843d313e8f851eda18385e19b95dabe02b5bce2a

Observation 13e18712-028d-4062-86a6-8106ca9f484a · inbound

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment cites this paper.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment UNITER: UNiversal Image-TExt Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.881258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.881258Z digest=sha256:5db88e502501b6560eddbd9a49ca55d0486a8c6acf6c8f1437c62b142ee2942b

Observation 45c0e4d9-8c5e-4266-a334-b1f1403cb4a7 · inbound

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments cites this paper.

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments UNITER: UNiversal Image-TExt Representation Learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:44:57.760345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:40:33.082748Z digest=sha256:b1a829dfc530b2d015e80f998d2eb0c200199c38a66dc93360d3d0306af175e1

Observation cbc61fdb-edcc-469f-b999-c395dde7d123 · inbound

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval cites this paper.

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval UNITER: UNiversal Image-TExt Representation Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:12:08.524191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:12:08.524191Z digest=sha256:511f5b1d7d5d0b462c7db4aed894c58e8d7c6c727389c55f6b3995f52e9d67cd