Pith. sign in

Paper Citation Record · LEDGER

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2511.01390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.01390 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:26:34.931761Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13e18712-028d-4062-86a6-8106ca9f484a · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment UNITER: UNiversal Image-TExt Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.881258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.881258Z digest=sha256:ad61efe19ba9350c265d02c967460ffc21567e5932eabd80b7b6dd416e4efa0c

Observation e8e8c1f8-b211-4a09-9b66-1dcaaccc404e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.890073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.890073Z digest=sha256:ad5e577ae009002970540d0afd93c58d8ade4d11886b4f44d67d6ed2f83b7394

Observation ef988689-8f4a-486f-b352-e90a489da990 · outbound

This paper cites Decoupled Weight Decay Regularization.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Decoupled Weight Decay Regularization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.907274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.907274Z digest=sha256:51dc62c2bfc1d7da06953fae9c09b1b1e7b5939d7c5b1ca45414d8fd2f5c334c

Observation 9ba09913-986d-473a-bb3a-c946c2025e09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Learning transferable visual models from natural language supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.915761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.915761Z digest=sha256:14bc8e2c950017180f3511b9a7e8a26554347ec91322e216782ac1b606210320

Observation e41aaab0-26be-4162-815c-c83f2c3eb5ed · outbound

This paper cites Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Table 4: The comparisons of image-text retrieval for SEPS-Vit and SEPS-Swin with different selec- tion ratioρon Flicker30K

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.927587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.927587Z digest=sha256:a89be1a5a649c19db6738e36ba07b8afb9c36b5306a53e4d456056f8d7fe18f2

Observation 4e09f437-9490-4a72-a487-2cda4586577e · outbound

This paper cites an unresolved cited work.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.931761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.931761Z digest=sha256:a292c388a848cf4524d49aab696f8ea5947e621dcfe7bb0ce5ed1a7122f37953

Observation 56fa954e-cab7-4085-88b5-3f7fa5188f2f · outbound

This paper cites Image-question-answer synergistic network for visual dialog.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Image-question-answer synergistic network for visual dialog

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.898696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.898696Z digest=sha256:82894ae91cab4c21513f302a91d8ef5af952302220061605bb3fb63161ea965b

Observation 627a55b4-1432-4506-bed7-97b5580d9b60 · outbound

This paper cites The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.911429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.911429Z digest=sha256:4e8c12dc581d97d7f294db240efc3a5b3f4bd1d3ccff76cb224469a59eb51145

Observation 14faa27c-3a61-4dfa-9a5c-93b367c8ba76 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.894136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.894136Z digest=sha256:f789c5807afd05d2fbeef91e781c24fd30fb2888f833157635a085355bbbe28f

Observation b6643d60-1288-4c2e-b079-7bd486bcfd1f · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.886208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.886208Z digest=sha256:033f1b4eb89346aeb4ab98029cb5b4934f7649158941eda36d787f74ac1e13e9

Observation 4fa0c16c-36bd-49c5-9b8b-91f1bd5ba792 · outbound

This paper cites SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.919618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.919618Z digest=sha256:11ed040cc446b7f3cec5dee4f5cf267d16dcc0d9a9ac0782b0a3134448b97320

Observation 2cd20453-59c6-4e10-af28-860e8b47f367 · outbound

This paper cites Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.902867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.902867Z digest=sha256:a37ae47307e9fe6d8a0de26afac7d9197049c1469103f4ff5507643cde1b1230

Observation 1efe51f5-7dbe-4d7b-9141-9e9ba65ec5e8 · outbound

This paper cites In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap.

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment In stark contrast, SEPS not only significantly surpasses all prior fine-grained methods but also successfully bridges this performance gap

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T00:26:34.923686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:26:34.923686Z digest=sha256:a1488873d5e9a952e12990991fde2171c5073b633e95923c98ed23603b3eddd9

Pith citing papers

No inbound Pith citation observations are available.