Pith. sign in

Paper Citation Record · LEDGER

FG-CLIP: Fine-Grained Visual and Textual Alignment

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.05071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05071 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:54:59.478902Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.758647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa57dabe-d4a1-4abf-9dbc-0aa92135ffcd · inbound

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning cites this paper.

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:59.478902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:59.478902Z digest=sha256:d5b98b6db5a8e340a06742adf5756cbbfb4d60676b13f306799a043ad3aa76c9

Observation 5da17b1c-2b8b-44f0-98a6-3cd8ef72f4e1 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.099238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:58.099238Z digest=sha256:414eddcebcec41eb84a78c4e3fd3de370406c091573280a38ee2daab4ce14eb3

Observation 4611edf2-8eb8-40cc-b249-c6b37710c0e2 · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.615646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:4ef22a11d5dc1dba94db12313b12acf453297131a4e45e9d7e617c4ff1bdeaa9

Observation d670af6b-6e6d-4657-919d-66937fa6d802 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.287769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:c6de64e4ddb930a1dc4dff5c753995572d949b917a0c7004ea90a0104306cecc

Observation 5286bb90-ce9e-4975-8e5e-ddd122d783b9 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.553656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:0e62271fbefa3a3f4625ccfec2f17fc9e41cb9d97721d107ee66cc0dc7e2effb

Observation c175fb72-8b60-4568-847c-979d9611ee67 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.036252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.036252Z digest=sha256:ac3ab672f54c3a40ece1fbb43e87865ef6c28bd5951b1e6e454c1aaad62be342

Observation 7373a016-8256-41ca-82f0-b07f3affd1e9 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.903003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:08b4fcef1a7af5bdd901637050e4996e53b2d63a32091a53506db6330fd771eb

Observation 6bac1155-41c3-4421-80a9-d973d6a95d1c · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.259432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:6687daa48da5b1f564020e0f6ff34fad4230790c1eb5ab77e083e70330dec15b

Observation 9f7720e3-99ec-4222-b6b1-9508dbd3ab9d · inbound

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce cites this paper.

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.736470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T00:55:56.146885Z digest=sha256:01cc06a30f2c97c3080d9b335783000ec9df5c56342fc9e0b8860ce6b05d87cc

Observation fc4fc96e-58b6-4b8e-ad78-3a6af6b69bcd · inbound

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations cites this paper.

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:09.981957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:33:24.911067Z digest=sha256:21de6c4b62d954d138dec58e1fd52c33341e4ce2d49f66fcd85f633e513e075b

Observation d7bc418a-5b55-4a14-8883-44847e771448 · inbound

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment cites this paper.

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.757322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:38:09.299209Z digest=sha256:ed20f0fc4e4a983b8649a6825f165e03e968c6e6f0aaa355cbc417832b5e8306

Observation d593473c-b77d-4a0e-ac67-f2dd3a9717c1 · inbound

L2P: Unlocking Latent Potential for Pixel Generation cites this paper.

L2P: Unlocking Latent Potential for Pixel Generation FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.727524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:12:28.181595Z digest=sha256:393dbf6b5510a37f4492b43fafe153c5990b47e1798234a2329aecbfac9312fa

Observation b4d4030b-f343-42a7-947b-d622d4549ec1 · inbound

CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection cites this paper.

CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.582680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:13:34.998844Z digest=sha256:ac5f14a2cc2af09bc3478e4656740480ef0fcf3277afcb9fd994a34d3976fd1c

Observation b30c20c8-4764-47fd-9777-738b298868d9 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:40.789072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:14:35.109298Z digest=sha256:fde924a93b7f48b16ed4300c4a4f022ca3cbdb5359bf52bc1a046f8eba01d6f5

Observation a44653e4-c4fa-4e70-b2de-5d8a0cb884d9 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.629755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:00675a523cc7ce395064430cd749ed66b94423f2926d7198f3b8ea8a82653725

Observation 04bb5bc8-c02d-4539-9e83-84470608a121 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.760118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:024b6d86a079db4f35f8639a5a41911520211b90fe892f9a2c2e85f2e92601d4

Observation 85a38dce-0093-4b5c-9272-47eb371fb75e · inbound

InstanceControl: Controllable Complex Image Generation without Instance Labeling cites this paper.

InstanceControl: Controllable Complex Image Generation without Instance Labeling FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.109676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:37:41.030752Z digest=sha256:9f8c4a8486a1a44d913f3a2ff45df49e1682add88536e32e0be435406b978c07

Observation 8addccce-036f-48d6-98bb-def375b1dbfc · inbound

DialogueVPR: Towards Conversational Visual Place Recognition cites this paper.

DialogueVPR: Towards Conversational Visual Place Recognition FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T14:39:42.933706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:39:42.933706Z digest=sha256:e13a88bc0ca1a39ef9affd707e319abe948f24651a72f06625f245bd030a95da