Pith. sign in

Paper Citation Record · LEDGER

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

As of 18 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2507.10095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10095 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:58.308176Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:48:58.248509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:48:59.054054Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85b82364-f8a0-4410-8018-94e29995b356 · outbound

This paper cites ComAlign: Compositional Alignment in Vision-Language Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ComAlign: Compositional Alignment in Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:47:58.985842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:50.844567Z digest=sha256:3e8a9a9c6373a853b7c9ae3922b21bb5903bf898e9a87cfce1848a719aaf979e

Observation 5eb06bef-2fc3-4f30-9434-d7eee1608165 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:50.931195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:50.931195Z digest=sha256:89b964504b582eed611effc92ea29f1781a3bcdf15867284eb218ea4321dfac4

Observation afd3135c-ea32-4aeb-bf47-f167dbe7863a · outbound

This paper cites InternLM2 Technical Report.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.023553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.023553Z digest=sha256:52c3664a4c05fd17feee9489ff170342ad81649d03dcd0d6b5c71bfa8ecfd15b

Observation a0393b45-96bd-4029-b95a-7c94ffa108bf · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.123031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.123031Z digest=sha256:572ac2482610d50eeb5241b66502784e8f9969dadce25ea13ea1118270e500f5

Observation 70c2c9f4-f23b-4a93-8bef-33bfa0a088ba · outbound

This paper cites PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.190600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.190600Z digest=sha256:1ebbac54df72c6f733aabdf8939f3e1c000339bf5eedbce7f8b790aadd718a59

Observation ed9adfa0-7708-43dd-8730-9ff4da8f548a · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.300206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.300206Z digest=sha256:8f01892a29dd7c9ac1bed63ad2794930bd91702a9d00efd52a76b1a695dd041a

Observation 28858f04-a530-40a7-a1b3-995cb6c67809 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.375599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.375599Z digest=sha256:c2b73f29f0d5e84b880c3fb4b3fc07aa92dcd9d5f6117377b9d899f8a47a075e

Observation 938540eb-a592-4be7-8f5a-1fcd2866819a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.456033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.456033Z digest=sha256:99986e2565685f637811dddd87dcb4c70d393a97593e9546b6c6c25de1f6d8b1

Observation 7d540b81-e646-4792-a26d-645d458057f3 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.529823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.529823Z digest=sha256:64227cbbea2afb311f3e1bf2b59e462c9233af24e5219675621165260238d98b

Observation ade1a9c4-2ce5-4298-96df-6fdf3230edad · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.616122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.616122Z digest=sha256:e1ccecbd9a26b8ee75b2058447242f084b09c94ff7988c295f252c7a85a63e05

Observation 4faabd50-9abd-4a31-b138-b46392c1996c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.691092Z digest=sha256:b9624e359069445e8d46d86c38df591f2c6cfc27f1b23f2fadb110960ea5976d

Observation f8e15959-695a-42b4-bd32-9668e0d48eb0 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.757234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.757234Z digest=sha256:7e9dd04ad82f01dce53aec3a7e9f227ca704018b13cb5ad409f00d7af2b54440

Observation 899646ea-d9f0-42c5-b499-564324a4c96f · outbound

This paper cites Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.836489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.836489Z digest=sha256:137e298ab973189e16a96ac2b30f26c0567c509104d2cf520f8ce095b8654201

Observation a088e4c0-c9d3-4ddb-ac43-352e9f520b67 · outbound

This paper cites Improving clip training with language rewrites.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improving clip training with language rewrites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.895009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.895009Z digest=sha256:2ff0ff8790ce9342a668130602f2a5b6e99e00d1c85fef82b2f336362ba222a6

Observation dea49971-4fc1-4e00-81e8-0d2aecc88b25 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.976190Z digest=sha256:ad33520071df9fc6e31814b5be35bfb6d0e7c50bf4dbf248715addb7d7439bdb

Observation c0fb7b0c-7aae-4444-9522-df10b1fbf046 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.096635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.096635Z digest=sha256:9e376a092f40155b8a5cad0c8b30c6ed023bf584088603d2d1833dff1fd7ad5e

Observation f37346fa-c5ac-44bc-a306-54ab2c5f6b8a · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.167649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.167649Z digest=sha256:7e37bbbe0e451355b6c1da15791602b12d84a6afaa1c5e3ee0984edcbadbef0a

Observation 56d02a32-66c5-4526-996d-8626deded757 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.241373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.241373Z digest=sha256:1821f414b3c78039f3a0fa54cd29192861885e47b298da2defc442139c410c89

Observation 3bd88bcf-110d-4fe7-a18b-43e7ba3f1b94 · outbound

This paper cites Masked autoencoders are scalable vision learners.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.301323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.301323Z digest=sha256:1d1fec5025b2ecb7d6c5d4619819c55ee995551fea4ace03595f4c69bc8a587c

Observation 0993e728-622f-42ee-8eff-7ab51f8dbd05 · outbound

This paper cites Natural adversarial examples.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Natural adversarial examples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.382258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.382258Z digest=sha256:8ff915e8591458b8ca4a4d131ff792125d5a62be9e26b8044055595fb7d565b7

Observation 175916ac-0f86-4e73-a6b0-5db895b865da · outbound

This paper cites Dynamicid: Zero-shot multi-id image personalization with flexible facial editability.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dynamicid: Zero-shot multi-id image personalization with flexible facial editability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.468856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.468856Z digest=sha256:66b1ffe644cb83e9921bdadeb283c5ed4124371febce6e5416fdfea1b3d7bc87

Observation eb2da0f2-6969-471f-a46e-15dc5b7c6095 · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.522994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.522994Z digest=sha256:3ce6d56650a435acb59deb250f59fe2b2ba19613e3260fbd95e3159cdc75b102

Observation f184e6ab-a845-4049-ba47-7fba77f79468 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling up visual and vision-language representation learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.606217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.606217Z digest=sha256:31bf5fbc48e4a6c01e24ff785cb8865c1044dae97210ddf2d634e87e23659b96

Observation ea6a9366-3a8d-4047-b5c8-45ebcf0c6d39 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.671434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.671434Z digest=sha256:19882e5900c6f6a16af6bd838b61ef6b90154a093d1fad2b595da0a62cbfe7d4

Observation f8bbd309-d12f-4dfc-8d96-8c87a5b26b65 · outbound

This paper cites Krizhevsky and G.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Krizhevsky and G

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.755130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.755130Z digest=sha256:7bfa0a05dc2bc37a6377a5533b7e8716175e0bc1d38d7a91e5778dae9440e3fd

Observation d5a3b105-2cab-43fb-9598-5497516bbd4a · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Veclip: Improving clip training via visual-enriched captions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.863818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.863818Z digest=sha256:cf20508a11dabe17decd590353b9d87cab9b0c61f242133c9e59a32dbaa340f3

Observation 9707efd1-86d7-446f-9c2d-266ad069bae7 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.943918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.943918Z digest=sha256:f29d153e06c241f0ae62c12e27a672890143b5f133bc19020d2e1a00160e7734

Observation 1bf5e39d-976d-4250-895c-595d092070c6 · outbound

This paper cites Language-driven Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Language-driven Semantic Segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.066777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.066777Z digest=sha256:ad204f368820d84cf13518fd4019b0050a7367684df88be8236567c36534a482

Observation df3ab08e-13fe-4c2e-9915-ff13b3447768 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.169799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.169799Z digest=sha256:63719a125d2d29f663d48c392b3f2299707ab177210d9e50fecd1ce351965ed1

Observation 38902c7f-4cd6-4b7a-88ad-60d35621164e · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.279551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.279551Z digest=sha256:2634ed2592fa66fb242a9deb01d232008876225b5488f4c58c3892a07da7ac6b

Observation 689f7079-8927-4709-ae38-99d07585b202 · outbound

This paper cites Grounded language- image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.318476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:53.358274Z digest=sha256:e94dcb243895a9aa8c69f39d965fc23504cecb588d730b22faa65bbb6845961d

Observation 8b5e591e-c7e2-422f-be41-a434ecd1d60a · outbound

This paper cites Scaling language-image pre-training via masking.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling language-image pre-training via masking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.302907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:53.439117Z digest=sha256:8c1d4b94755e20d7df152e2b22790f8c1695321bb379a03b1cd8a869e4da3444

Observation f75d187f-ca3b-4a50-bb8d-30094aacadc0 · outbound

This paper cites Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.532210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.532210Z digest=sha256:32e2c05823d4204068d25d3952ea5e56d2f4370e90a72a551a1126f3635af3c0

Observation bd69c4e9-a10d-4040-9c09-556be97cd576 · outbound

This paper cites Improved baselines with visual instruction tuning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.632235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.632235Z digest=sha256:a119683ad8f61b0fcbad6d9baa67a55216fda4229ccadba035d9e3fec4b587b1

Observation ce4e1047-352d-4c57-baa3-160ebe7909a5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.276811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:53.688289Z digest=sha256:b3ca44541bd954f412d242433156d51d8a40bf224d1199c9e91da1bc140c4f5d

Observation 87c2de30-0b79-43e9-b44a-1e6ae5637304 · outbound

This paper cites Decoupled Weight Decay Regularization.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Decoupled Weight Decay Regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.754810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.754810Z digest=sha256:e9723485093aace874d7d202dbdf291054be9158070824e5f387918d4bce36c7

Observation dfe63ba5-7b7c-4215-a763-0e7e4f8ec317 · outbound

This paper cites Open vocabulary semantic segmentation with patch aligned con- trastive learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open vocabulary semantic segmentation with patch aligned con- trastive learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.261207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:53.837321Z digest=sha256:38a8b1a319fbdae47586ea014b8943c4e3ee93a124adbf35047702cf173f0644

Observation c1f8caa5-333f-45ba-b831-33fdf3fa9b2c · outbound

This paper cites TULIP: Token-length Upgraded CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text TULIP: Token-length Upgraded CLIP

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.923146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.923146Z digest=sha256:dbf363d0c29448d52fbac7500a652fcbc0466bfb73cd1961c19b667254e1317c

Observation 029cca05-7d20-4c51-87fd-21dbe3b0d59a · outbound

This paper cites Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.246042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.028385Z digest=sha256:763fbfab9f648014964ed4276a55fca1dfeefd126d57c9f084b65afd7d80b932

Observation dc291d73-dc63-47b4-ade4-79c070d4be35 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Im2text: Describing images using 1 million captioned photographs

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.230216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.123427Z digest=sha256:ff7671f2a1ee1a832fe1be6d2a6ca7e4e10c38f3d599897f1a6792ba2b9d1d79

Observation 6f0d734b-6310-4ce9-90ef-036f7181759c · outbound

This paper cites Scalable diffusion models with transformers.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:54.242148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:54.242148Z digest=sha256:7e2ff98ebef615d4846cab60faa57c82c3f0a8652d3466834c4ed682e5dbf67d

Observation cbb2f623-2e56-4221-8d1b-a3ef970c5bf9 · outbound

This paper cites Bizgen: Advancing article-level visual text rendering for info- graphics generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Bizgen: Advancing article-level visual text rendering for info- graphics generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.200175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.384756Z digest=sha256:02d012d985fb46bda1abcb477c61989226f4a3ac3f6e2e8891417603bc872e27

Observation c7914707-cbcc-420d-9889-39d67cf425e5 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.185478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.522718Z digest=sha256:9b64608a1c9f9744758aebecc7e5b61bdaab6ca62e489d55b052c02ecfade650

Observation 28f7ae4a-c9b4-4fb9-8716-61328c3dfd68 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.645406Z digest=sha256:92004165a7370b36e83841117a9efcf142b4cabeedbe9d7d9e886e83819bc168

Observation bab91242-e00e-4b59-90fb-3b242f9cda8a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.155853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.728047Z digest=sha256:26376a2a702ce4a1155224b245d0572e94a6ec2fc9d0d8d07f9ff4e389d93a81

Observation 727ad07b-659c-488b-9d1e-ec414cfc7994 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.141043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.847562Z digest=sha256:45e6dbc88f2218ad4e18678bf54e54d5caa18485663f4d672f162a8a02f1923e

Observation 0ff8bf5a-e4fa-4a0c-a26d-8c423df35567 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.124884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:54.951984Z digest=sha256:9059e13f1d8ee145834205a7c13a8f2160c3169480b628794d86b2f22a6982f0

Observation 245d499d-fe53-48d0-bb61-499da5580eb4 · outbound

This paper cites Umg-clip: A unified multi-granularity vision generalist for open-world understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Umg-clip: A unified multi-granularity vision generalist for open-world understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:55.094107Z digest=sha256:4f15e4d8e52f59479602c7626ec57f8d2d7a7848c9fd34641ae6ec7067af532b

Observation 3d6d531e-fd2d-4f98-a773-afb55277d8dd · outbound

This paper cites Localizing Objects with Self-Supervised Transformers and no Labels.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Localizing Objects with Self-Supervised Transformers and no Labels

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.183779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.183779Z digest=sha256:4b0d9cbd5bbba9511ac6cff73a41c7235f53b480bdecaafb5f57413ec2c28ed8

Observation bc494d9b-2f1e-44c9-b2e0-76e80a0af7f5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.324952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.324952Z digest=sha256:d76dc4232a92f030cc9abe04d3175f5583e8c9cef440d7a96de30c9175fcf17f

Observation 4e448dc6-6881-4b5d-a333-17c921c4b40e · outbound

This paper cites Yfcc100m: The new data in multimedia research.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Yfcc100m: The new data in multimedia research

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.095838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:55.486925Z digest=sha256:255de1363977e084e14e76306efb782cbf9c4b480a6285dd330902d0a959abe1

Observation 65efc709-fe4d-4d36-9117-1bc24eb43b2e · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.080064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:55.616713Z digest=sha256:11bc57580d8638d8eb418e897f2d5e6e06c52323d8768af16ee3d6fee7cc173b

Observation b3c68827-f803-439e-90b3-2966eeb41bdd · outbound

This paper cites Position-guided text prompt for vision-language pre- training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Position-guided text prompt for vision-language pre- training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.061108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:55.737623Z digest=sha256:35c7553a1b2630182c12c8e6929fddccef7196a24ccb24e7a7f089925961f0f9

Observation 7e086497-050a-4e97-acf6-b36067dc6e10 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.044173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:55.855959Z digest=sha256:960ca9eb83070332bdcebc84b12be1330dc0031d14f7ba35bd2b7edda103a91e

Observation 0a439027-8353-4b1c-981a-0ec2b4a4eb96 · outbound

This paper cites LoTLIP: Improving Language-Image Pre-training for Long Text Understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.987654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.987654Z digest=sha256:d5218f267f20adff798ee852533a3d2df61aedbe2f24ae88bc142b667dd9231a

Observation e3c8eee6-d1a6-4c8e-8fae-a6925cb17173 · outbound

This paper cites FLAIR: VLM with Fine-grained Language-informed Image Representations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FLAIR: VLM with Fine-grained Language-informed Image Representations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.051883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.051883Z digest=sha256:ac72f796c4ded682976906fc53a855e699729c58b43ae58972b88cb380615f6d

Observation ab5e76e4-5cb3-4201-8efd-22c5c53a1fb7 · outbound

This paper cites Demystifying CLIP Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Demystifying CLIP Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.114498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.114498Z digest=sha256:f5e3224030b95f6b43f90d5e80038414b0a5f9430b4903d7e633acd1878d8177

Observation bc37e39e-0277-4cf0-8044-0384f944813d · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:5193e20cd1502b9f58da1cdaddf7d4ceb708be38fdb44ec6cef68db506ad8834

Observation f1bee241-eed6-4c4f-a81d-9147de762f80 · outbound

This paper cites Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.027925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.218338Z digest=sha256:083ae0cc50de54fe78a27d05f4f639f5147b3945e6773a8dacab866d436e1a34

Observation 5ddd13ad-b507-4f8b-a92e-e35421b82a72 · outbound

This paper cites Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.009948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.315764Z digest=sha256:6aee0ef3374b8345b6444e91812175f04e5fe0a4bcc137a1d71e44900f7a487d

Observation ca00cd65-efd2-4abd-a442-23a2eee6796e · outbound

This paper cites Capsfu- sion: Rethinking image-text data at scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Capsfu- sion: Rethinking image-text data at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.992184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.396155Z digest=sha256:fca5cf639c00e030e86782b1d8d2d59cf7fd34ee6b0ae184e9ee8fcee43c4185

Observation b7b5ceb2-96e7-4128-9bb4-28c5d0c7d0ec · outbound

This paper cites Sigmoid loss for language image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.972522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.438155Z digest=sha256:b68d782f1abc3e43b63783f936b5815f3a34d9575e3d05b0eeb954f92f2cb896

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:c67fbe8ea791bd1ea2b6581fb498c16df7c8984a4d31506b321b7f86ed0d6bcb

Observation 9735214e-b1a4-4d5f-ab04-97a552d3d2be · outbound

This paper cites Vision-language models for vision tasks: A survey.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vision-language models for vision tasks: A survey

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.618386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.618386Z digest=sha256:f36bec0ad29b71c16a1c9d4c4b73c87765d0b5737218ab76b092eccc28a10200

Observation 21800da7-651c-4c08-bc12-c727d5b80f70 · outbound

This paper cites Exploring regional clues in clip for zero-shot semantic seg- mentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Exploring regional clues in clip for zero-shot semantic seg- mentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.936071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.706784Z digest=sha256:5f254a99429efbc9e4914710c52d9463f045649761f37ae6dbef788958472d0c

Observation 5421ac72-bac2-442c-a643-c9411cda6979 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dreamlip: Language- image pre-training with long captions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.914531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.791346Z digest=sha256:86fd9eb11d8b4154f2b59dd521643c27e3a71cefc109e3753b1799872fc2fa51

Observation 462a7847-6e62-4390-99db-6e3b41da792b · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot semantic segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Zegclip: Towards adapting clip for zero-shot semantic segmentation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.899086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.860708Z digest=sha256:4a97917e9cb054f5387df86f4fa52b7d2cf01335b5e5d4ac60b42a4266ed05dd

Observation 3f7bd7fd-85a6-43f1-882a-8fb7ae5d882a · outbound

This paper cites During the re-caption process, samples are randomly taken from the following 20 prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text During the re-caption process, samples are randomly taken from the following 20 prompts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.880236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:56.940555Z digest=sha256:7bfd5e706ff5d5be09a00c9bb25ab87de9b6240279873c191f089637b8bea82d

Observation b27e8572-25d7-417c-a600-bee0d405dda4 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.864988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.004886Z digest=sha256:0fde03c3f3cfc58620b09fbff2877a881e7394a33044d2f1764a6f5ac446edcd

Observation 2979e6db-0f6c-4255-b2db-473a56218dc8 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.072569Z digest=sha256:1057f93ea2f6e4b5657f6033cd23de482a6f21a0dbd61941c2a801d1546adfaf

Observation 630a3f3a-739a-479c-b434-737331f99d67 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.834782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.135153Z digest=sha256:fb8b7376e99706a629a640be829b3a77bc4f76bb9554c0eb07c76b9ee508d91e

Observation 292b5fe0-9da0-450f-a811-30b2b2ae16f9 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.815510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.256816Z digest=sha256:62c51c01fa36d596a3af75f43cfa5d8503b036ce9171287f3ba0f566b897ab64

Observation 4399dc0f-e3cb-4c85-9f93-31825f001afe · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.798569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.327849Z digest=sha256:72da6749ac988a9355d9e6a34632f1d3901ea8a245e51bf688e8768d0fbd401e

Observation 5fb6db72-521d-4eea-ba5d-8d15d5f7b6f5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.779688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.390660Z digest=sha256:db784d91bc0c80b93a75e87396f4697be14511be215c4d3bf863f981ec9c1682

Observation 16ced39f-51fd-4ba1-bc08-1f1ec6ad927e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.763152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.496477Z digest=sha256:fac88790966ec3109966bed71afb5341082177bd2cf3ada678d9e2be7c94d266

Observation a5479ea5-f9f6-4c66-8327-1e0cbdf7ad6b · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.736635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.556587Z digest=sha256:d29b494ab7add9bd340f216b8f3db1df52beeee7b3bdcfb093ed216f37a46240

Observation 334b0f3c-161e-4fa7-9fd9-51f853f72bb5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.712612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.621546Z digest=sha256:d06d97916699345800c532db7014e30f9f45c27140fa433807d3774e4b796e50

Observation 2abaedfc-af62-469a-8c19-7697cfaef980 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.691201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.723436Z digest=sha256:88c4158f0d0bd402080ee3edd13db3cbf45bf7f73395a357c4eb8f66c44971b0

Observation 8510c5d8-f4c5-495d-bbdf-be6ca6436edb · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.673643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.764626Z digest=sha256:1d83576dca2f6155d0f6149350016d92f79b2a5f08e4dc42b636ecfde1fcec36

Observation fd9f8879-87e2-4a54-84b5-59f3da917bf5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.652297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.828019Z digest=sha256:af1f5836d125646a93c7f4384c6210337aebb1db4b2ecf7ab7a3eaddeee7bdd7

Observation 22ced992-f3c1-4445-9f28-1447279b278e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.629924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.895073Z digest=sha256:b81c4c37a2be30b36c63c21f84ae7eac70e9331f911ea70e66bbcd6e20963eb0

Observation 9818ae67-df68-4042-a7d5-562a803eab9d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.600287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:57.960361Z digest=sha256:860fdb0eb3629da3888847f213cf5b38e8f95d1ccb56423580a74693d9583b3e

Observation ac6ce2f1-36ab-44f4-808a-f4d137033c05 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.578432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.024087Z digest=sha256:577bcd07a8fa2a98d2bd5e0f2c345b0d5f833e116c6bcae2b5df0c63914a1004

Observation d9984249-4631-4ce7-813a-f1ef90a0a32f · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.556838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.065219Z digest=sha256:2747d6d89c4c41f1611d5a79e600602e144cf0bba71171560453f2d626338fb6

Observation 9a450b0f-a173-4834-a6fb-61b3e8b55c1d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.536723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.176510Z digest=sha256:b0dbe670330d9f4cad60c72971e0f74a6a32248f144045b6d7f8f57a7f548ffe

Observation dd41036f-1848-4dae-97e2-edb3d32b46d5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.521128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.244879Z digest=sha256:b3cf4184e5938cb03392e5d5a378a2d3988886c060ce0fc52b1477e30fdaf32f

Observation cc697d6f-033e-4ba5-b1f3-f299809ead8c · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.505202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.274042Z digest=sha256:844b886a87c3bcfc5380a6e79f14e2e6ac7a7272aff3686dfdc475f2bade1f87

Observation c6a95066-2bf9-4323-96c3-827cc939b0a7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.487803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.279069Z digest=sha256:7e87a57caafd416e4b9fdab8bbf28f6a375dd45c9ffd5d7ded6081be2ff9a490

Observation 7318b5a6-9db1-4156-a524-9e236f79ca00 · outbound

This paper cites We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.469236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.284468Z digest=sha256:2a2755226ebe8e630196b346208308343161f686734c9801bee80e578ab73e85

Observation dd714861-6b4d-4d9c-b50a-12f83254acce · outbound

This paper cites 5th of October, there’s a significant event highlighted in blue - the launch.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 5th of October, there’s a significant event highlighted in blue - the launch

Reference 90

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:48:00.250320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.289503Z digest=sha256:fddd1bc08e7a736f392b317ab0e7de49f16c0355c28e6bb40a882d77885c293c

Observation f0a4420a-820d-4fcf-864c-64efaf7c91f2 · outbound

This paper cites Shared Prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shared Prompts

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.972954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.294865Z digest=sha256:40fc5186211997a7de92d4453f229ff42293db088ef2e9f271f4a9a0489edbc2

Observation a91ca14b-6d26-4954-9d66-065ab404ec9f · outbound

This paper cites 8, the regional prompts obtain stronger responses in the corresponding local patches.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 8, the regional prompts obtain stronger responses in the corresponding local patches

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.679889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.298832Z digest=sha256:488543be33f132fc8e74f68da3b40dd42508410df84b5947d3af2ea6bcb6f590

Observation b917b074-82b1-4f5e-9acf-087729fe5ba7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:47:59.460818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.302953Z digest=sha256:7e1f0e5e5e255adaa727e079534799da7341b307c326e2b4928dc2080563b14a

Observation a663cbe5-11fc-4b8c-be9f-c887bd864bbe · outbound

This paper cites We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.238484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:47:58.308176Z digest=sha256:aeb3c84467a3d6285ec215e4bed418419424477d6d17f54d82ce79b769adb68f

Pith citing papers

Observation 660bdf26-2245-4bdf-8e7e-b8d0ad692313 · inbound

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis cites this paper.

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:48:59.127431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T05:48:58.248509Z digest=sha256:fabbbbdc4f675d2540be5535137844c0553f7e0c70a3df696cd6ff7e07a23490