Pith. sign in

Paper Citation Record · LEDGER

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2507.10095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10095 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:58.308176Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:48:58.248509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:48:59.054054Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85b82364-f8a0-4410-8018-94e29995b356 · outbound

This paper cites ComAlign: Compositional Alignment in Vision-Language Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ComAlign: Compositional Alignment in Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:47:58.985842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:50.844567Z digest=sha256:051f498a7425e5107293322c40b9ca0f0b7c25ae97dda50d29e34c9bf5df0216

Observation 5eb06bef-2fc3-4f30-9434-d7eee1608165 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:50.931195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:50.931195Z digest=sha256:9ba1e9942d1fb44c1dff4f72b9d89817b58587d4fec626bacc269fb45541e50c

Observation afd3135c-ea32-4aeb-bf47-f167dbe7863a · outbound

This paper cites InternLM2 Technical Report.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.023553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.023553Z digest=sha256:cf92bb00d15e25e9bfc0554423c47c13e93593092bc145a6a238cff590900108

Observation a0393b45-96bd-4029-b95a-7c94ffa108bf · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.123031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.123031Z digest=sha256:749bcde078d8acf25f235d9ece8f3b1794d1ab2b9b68e6218cf9347d5704a436

Observation 70c2c9f4-f23b-4a93-8bef-33bfa0a088ba · outbound

This paper cites PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.190600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.190600Z digest=sha256:0edaba478f6f357b80b28af9375a923af3cba8c845c6a25c639b25716b1c19bf

Observation ed9adfa0-7708-43dd-8730-9ff4da8f548a · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.300206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.300206Z digest=sha256:981af4861a7d37552875687dd9d6ad22134c227a5faa9e4cf879cc2856acdbe3

Observation 28858f04-a530-40a7-a1b3-995cb6c67809 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.375599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.375599Z digest=sha256:9e582ba3b41ec314b020c014875974209daa25663be72734e4b124f43f7b3f5c

Observation 938540eb-a592-4be7-8f5a-1fcd2866819a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.456033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.456033Z digest=sha256:5dcfac4970a255acdae55a206cca7fc353ee9caebe0dc8bc473f1a5daefdacdc

Observation 7d540b81-e646-4792-a26d-645d458057f3 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.529823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.529823Z digest=sha256:08ba60a525d1e4e6b1c6fbf8ec5ebd274d9ca1e24b37909cb1bb06949eb52482

Observation ade1a9c4-2ce5-4298-96df-6fdf3230edad · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.616122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.616122Z digest=sha256:5d1ff100c5bb5e393a0d0d3f6d357ca5a455e53c27142ea556ed0102cd7b20e2

Observation 4faabd50-9abd-4a31-b138-b46392c1996c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.691092Z digest=sha256:d6b4d7078bf9ad69425c0566fc2e62298cb84e79e4db4ceb3f4ccef8534140e0

Observation f8e15959-695a-42b4-bd32-9668e0d48eb0 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.757234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.757234Z digest=sha256:9a2a8ff77c44514f745aba43a00fcf3a056b8de5cbc5ca3338f74fecfbbef0f8

Observation 899646ea-d9f0-42c5-b499-564324a4c96f · outbound

This paper cites Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.836489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.836489Z digest=sha256:de0ee3aa203c53ea314f952388ad5905279b1a275a8d96f15a280f75d86d454f

Observation a088e4c0-c9d3-4ddb-ac43-352e9f520b67 · outbound

This paper cites Improving clip training with language rewrites.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improving clip training with language rewrites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.895009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.895009Z digest=sha256:19f4ed3a2956414399d1f95e50eae06c2dd25ff014ff4fb1c89c3fc6dfe7fd2f

Observation dea49971-4fc1-4e00-81e8-0d2aecc88b25 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.976190Z digest=sha256:de19f8404a3057ec5edf05889086d494d57d9d9393091ca1a82298df68847e15

Observation c0fb7b0c-7aae-4444-9522-df10b1fbf046 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.096635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.096635Z digest=sha256:47dcfdf7482b6e8545305d48e62f9a9b138cd128fa93fbe4bb2d39b036ce8657

Observation f37346fa-c5ac-44bc-a306-54ab2c5f6b8a · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.167649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.167649Z digest=sha256:a79ad4eba0db8e552b0621fadad3880044fca61b8a0d009d86656705820955d1

Observation 56d02a32-66c5-4526-996d-8626deded757 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.241373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.241373Z digest=sha256:4144b3ba8019215bc63693c1d1aec2003a6da4e7ad7be38f46e9982382363184

Observation 3bd88bcf-110d-4fe7-a18b-43e7ba3f1b94 · outbound

This paper cites Masked autoencoders are scalable vision learners.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.301323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.301323Z digest=sha256:6227a262d1ea7ebd9dd6a78bb818fd3f00ec06e3bd1cdb408a09f35ae44d7667

Observation 0993e728-622f-42ee-8eff-7ab51f8dbd05 · outbound

This paper cites Natural adversarial examples.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Natural adversarial examples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.382258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.382258Z digest=sha256:f498cc468bed2c3309cecf39eb87e3ef845f22089956169ea27b9ceb96728a73

Observation 175916ac-0f86-4e73-a6b0-5db895b865da · outbound

This paper cites Dynamicid: Zero-shot multi-id image personalization with flexible facial editability.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dynamicid: Zero-shot multi-id image personalization with flexible facial editability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.468856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.468856Z digest=sha256:599e132d5957b3e56dadcb9a8f92ceb9dd0e248c43fe804f6ad27c8147e71c00

Observation eb2da0f2-6969-471f-a46e-15dc5b7c6095 · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.522994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.522994Z digest=sha256:e5d16e4f1df61df5ba383db16075d9ba37f0cb8e551eb55e0e1b85722132dcc1

Observation f184e6ab-a845-4049-ba47-7fba77f79468 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling up visual and vision-language representation learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.606217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.606217Z digest=sha256:01d943a9991dcc5be6b338f7015f81aca5fad71e064f7045a2d8ac5466ba85a4

Observation ea6a9366-3a8d-4047-b5c8-45ebcf0c6d39 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.671434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.671434Z digest=sha256:69acb040d45d755fd74fc19eac56817bd21c823bb4871aee007acc597d734c37

Observation f8bbd309-d12f-4dfc-8d96-8c87a5b26b65 · outbound

This paper cites Krizhevsky and G.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Krizhevsky and G

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.755130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.755130Z digest=sha256:6e7442446c8761889e55c73482ebabe1d83fbca681bf56fe67834212ab1f3b6f

Observation d5a3b105-2cab-43fb-9598-5497516bbd4a · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Veclip: Improving clip training via visual-enriched captions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.863818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.863818Z digest=sha256:382a5069f63689292c4f50a26fee1f64798a984d01bf59e61262c0abc9134c4e

Observation 9707efd1-86d7-446f-9c2d-266ad069bae7 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.943918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.943918Z digest=sha256:c7a7f1725c03edcf5a8303e5aa687f6d5bf4b43c8a4c78a4941a5a0a71ece853

Observation 1bf5e39d-976d-4250-895c-595d092070c6 · outbound

This paper cites Language-driven Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Language-driven Semantic Segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.066777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.066777Z digest=sha256:dc93e29b21cf4f3258a2281cb94b5926aca40807f2d13007f721259d86c61450

Observation df3ab08e-13fe-4c2e-9915-ff13b3447768 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.169799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.169799Z digest=sha256:18d1d836308b0bbab9c927a47452875e59236e3aa0f867cd9d29caaf3adac566

Observation 38902c7f-4cd6-4b7a-88ad-60d35621164e · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.279551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.279551Z digest=sha256:d36ce683c4bb445a35d03354b8f9643e50319d3b28308f020835ef77656dac87

Observation 689f7079-8927-4709-ae38-99d07585b202 · outbound

This paper cites Grounded language- image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.318476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:53.358274Z digest=sha256:7628c672ba762d5aa434f230fc308e788c644cdd047db37a6131f54953bae2e7

Observation 8b5e591e-c7e2-422f-be41-a434ecd1d60a · outbound

This paper cites Scaling language-image pre-training via masking.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling language-image pre-training via masking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.302907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:53.439117Z digest=sha256:76f7c62cb10485111e6282f6863059dac1dd236fded8e8d392b597f203762eb6

Observation f75d187f-ca3b-4a50-bb8d-30094aacadc0 · outbound

This paper cites Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.532210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.532210Z digest=sha256:e39b1e31d3907f3c1b2bac6a5d0b237ee2f959826868993ab882edb6a77f2ac9

Observation bd69c4e9-a10d-4040-9c09-556be97cd576 · outbound

This paper cites Improved baselines with visual instruction tuning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.632235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.632235Z digest=sha256:6e8bef6586ce4f4e500296da69f53c539f02059f2517c2f0d35989da56d8507c

Observation ce4e1047-352d-4c57-baa3-160ebe7909a5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.276811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:53.688289Z digest=sha256:0542c3ee7af71faca6d7660011c0fa8905cfc67d06934187eb9e1bfa77b5abf8

Observation 87c2de30-0b79-43e9-b44a-1e6ae5637304 · outbound

This paper cites Decoupled Weight Decay Regularization.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Decoupled Weight Decay Regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.754810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.754810Z digest=sha256:9ad6157221f785fe579195974613b8b27ccb2226e4d4513ac6c1d69a1e6bb520

Observation dfe63ba5-7b7c-4215-a763-0e7e4f8ec317 · outbound

This paper cites Open vocabulary semantic segmentation with patch aligned con- trastive learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open vocabulary semantic segmentation with patch aligned con- trastive learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.261207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:53.837321Z digest=sha256:72c60e004539bd952a4c2067b03e8fcbb5b4cf6e70f7760ce760e478de520494

Observation c1f8caa5-333f-45ba-b831-33fdf3fa9b2c · outbound

This paper cites TULIP: Token-length Upgraded CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text TULIP: Token-length Upgraded CLIP

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.923146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.923146Z digest=sha256:0921c28445372a7cc807ee3fa095d859142113c49c4b485a6101634104a6b3f8

Observation 029cca05-7d20-4c51-87fd-21dbe3b0d59a · outbound

This paper cites Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.246042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.028385Z digest=sha256:05a5d8e318379be4a97ca7a03cd932771f6d8fa27f14843b1daadcd4d80705d4

Observation dc291d73-dc63-47b4-ade4-79c070d4be35 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Im2text: Describing images using 1 million captioned photographs

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.230216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.123427Z digest=sha256:2cd6fe3600f34ea17efe47f11c4be8a5ba73681db8e0b59b2ec5b12d75e38d16

Observation 6f0d734b-6310-4ce9-90ef-036f7181759c · outbound

This paper cites Scalable diffusion models with transformers.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:54.242148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:54.242148Z digest=sha256:b94da56801d995b2cdae331225b26b497a74fdac61e8340a3a0f4475677afbd6

Observation cbb2f623-2e56-4221-8d1b-a3ef970c5bf9 · outbound

This paper cites Bizgen: Advancing article-level visual text rendering for info- graphics generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Bizgen: Advancing article-level visual text rendering for info- graphics generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.200175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.384756Z digest=sha256:821e9fc7eb4bf634958edc45e7b1a840c68ba39229a0a17981e78d2e23fecee4

Observation c7914707-cbcc-420d-9889-39d67cf425e5 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.185478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.522718Z digest=sha256:b3da62a0135f834f65b55ff03e14c816839e42ac1d5e18285842bb278f95fe1f

Observation 28f7ae4a-c9b4-4fb9-8716-61328c3dfd68 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.645406Z digest=sha256:e3f0f34e56e2e3a3c57fe0e466584f9d84fb826ccc3b244f267e545afe1beaea

Observation bab91242-e00e-4b59-90fb-3b242f9cda8a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.155853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.728047Z digest=sha256:2bc293707e1f36e976e1f2096f8ca4b7dd596092f84a757780da1bac8f96fa94

Observation 727ad07b-659c-488b-9d1e-ec414cfc7994 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.141043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.847562Z digest=sha256:9490952f5181cbe0030140a63972596922d0525d56fc92c00e324d63ec33ee9f

Observation 0ff8bf5a-e4fa-4a0c-a26d-8c423df35567 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.124884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:54.951984Z digest=sha256:e5d7968c0e217492df8b3b64ee4312d63cbdb255f7d383ec39967635b83f1889

Observation 245d499d-fe53-48d0-bb61-499da5580eb4 · outbound

This paper cites Umg-clip: A unified multi-granularity vision generalist for open-world understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Umg-clip: A unified multi-granularity vision generalist for open-world understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:55.094107Z digest=sha256:25759d0e853938fc1b7c01b27d8ef35a0c590cbcec423ff2f4d8c6ad013c4d51

Observation 3d6d531e-fd2d-4f98-a773-afb55277d8dd · outbound

This paper cites Localizing Objects with Self-Supervised Transformers and no Labels.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Localizing Objects with Self-Supervised Transformers and no Labels

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.183779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.183779Z digest=sha256:65a05e7c7878b747b9ba8c0d7d26babe7e86b6f7fce004ba299de96a78fd0b4b

Observation bc494d9b-2f1e-44c9-b2e0-76e80a0af7f5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.324952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.324952Z digest=sha256:44c6aca039eddebbf80a98003a9f531adf02a9c4be3b4aabb7a08e459a5255ef

Observation 4e448dc6-6881-4b5d-a333-17c921c4b40e · outbound

This paper cites Yfcc100m: The new data in multimedia research.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Yfcc100m: The new data in multimedia research

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.095838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:55.486925Z digest=sha256:706521e983ecde978d9ee82d74d96af7e231b2fcff70544e68336d751c53cc1c

Observation 65efc709-fe4d-4d36-9117-1bc24eb43b2e · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.080064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:55.616713Z digest=sha256:9cede2124a1119c385c86cd248eb41b7c6d3780ac1d2514d8169f0c9aed170f0

Observation b3c68827-f803-439e-90b3-2966eeb41bdd · outbound

This paper cites Position-guided text prompt for vision-language pre- training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Position-guided text prompt for vision-language pre- training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.061108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:55.737623Z digest=sha256:26e89d507dbd1836c35f6afb59d35ecac96f6bf79d1c773766fc4eb37ad60d43

Observation 7e086497-050a-4e97-acf6-b36067dc6e10 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.044173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:55.855959Z digest=sha256:34b1bd2d08539afab024f2013c356f0106de86e2dceb6a3bfff0eaff5543baa9

Observation 0a439027-8353-4b1c-981a-0ec2b4a4eb96 · outbound

This paper cites LoTLIP: Improving Language-Image Pre-training for Long Text Understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.987654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.987654Z digest=sha256:dbead63e8f14caa719c49fa64bba3375acc6abeb07106508c4da7b43dcb20500

Observation e3c8eee6-d1a6-4c8e-8fae-a6925cb17173 · outbound

This paper cites FLAIR: VLM with Fine-grained Language-informed Image Representations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FLAIR: VLM with Fine-grained Language-informed Image Representations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.051883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.051883Z digest=sha256:6c3c9c19e4820790b3c041d8df36b3843e630e02e82520adbf75a7095a1995c3

Observation ab5e76e4-5cb3-4201-8efd-22c5c53a1fb7 · outbound

This paper cites Demystifying CLIP Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Demystifying CLIP Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.114498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.114498Z digest=sha256:bfbde29f40a3703f7486a663e45261d0746cc10616607b81e91f03d66bc49fa6

Observation bc37e39e-0277-4cf0-8044-0384f944813d · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:702f10124ea4a0a0ea140ad520591f9fb88c818ce4b5c7e5850ceef31dffa0c0

Observation f1bee241-eed6-4c4f-a81d-9147de762f80 · outbound

This paper cites Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.027925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.218338Z digest=sha256:e42ad817d64b7474c38bad08674fa3abdbc506226d5004df932ea3c03a33d097

Observation 5ddd13ad-b507-4f8b-a92e-e35421b82a72 · outbound

This paper cites Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.009948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.315764Z digest=sha256:c0685bf13b92ebb135e24cc1567451dad55ec3413b406c8d86549df165118934

Observation ca00cd65-efd2-4abd-a442-23a2eee6796e · outbound

This paper cites Capsfu- sion: Rethinking image-text data at scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Capsfu- sion: Rethinking image-text data at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.992184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.396155Z digest=sha256:119dceec119b7432bd23dd47a5671aede6cf7f1c0a6822258322dc32a79c7785

Observation b7b5ceb2-96e7-4128-9bb4-28c5d0c7d0ec · outbound

This paper cites Sigmoid loss for language image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.972522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.438155Z digest=sha256:d091bd7cccbb2398a85e80a23f85c261caf50b8c9f2c634be6f940a29e4e5785

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:3a64916703f0511d92c713bb532cc0e3678c1161e296e9c8b05fe01bde96c622

Observation 9735214e-b1a4-4d5f-ab04-97a552d3d2be · outbound

This paper cites Vision-language models for vision tasks: A survey.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vision-language models for vision tasks: A survey

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.618386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.618386Z digest=sha256:0d9d0d99d8379ea93e7242d51a504dd19f7c1fd48c6dc38a925c2de2d0e954d8

Observation 21800da7-651c-4c08-bc12-c727d5b80f70 · outbound

This paper cites Exploring regional clues in clip for zero-shot semantic seg- mentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Exploring regional clues in clip for zero-shot semantic seg- mentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.936071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.706784Z digest=sha256:8586a036a090a196efb5ec501481800296565ebb392cc335f502b4b3beef45a5

Observation 5421ac72-bac2-442c-a643-c9411cda6979 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dreamlip: Language- image pre-training with long captions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.914531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.791346Z digest=sha256:bf0eae63f7c4fabd0d68e90f91abc855831120199cf639f9c86e48981ed648ad

Observation 462a7847-6e62-4390-99db-6e3b41da792b · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot semantic segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Zegclip: Towards adapting clip for zero-shot semantic segmentation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.899086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.860708Z digest=sha256:afcef0575b7f87c29de7b472a1b8a6a2e8d84686f2f4b8bad2f1c2a970ae4e09

Observation 3f7bd7fd-85a6-43f1-882a-8fb7ae5d882a · outbound

This paper cites During the re-caption process, samples are randomly taken from the following 20 prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text During the re-caption process, samples are randomly taken from the following 20 prompts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.880236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:56.940555Z digest=sha256:6362c4cd85159b666e954b0f68b84b388448769705115e6aaf35945986a2e716

Observation b27e8572-25d7-417c-a600-bee0d405dda4 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.864988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.004886Z digest=sha256:285891a785a6941e26126772a1ab712e8e2e220d49092b70980d3cf2ded16192

Observation 2979e6db-0f6c-4255-b2db-473a56218dc8 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.072569Z digest=sha256:36994466aec9eac3ee4e725f6f0a94501aac0db8aea430cffce07ede130281af

Observation 630a3f3a-739a-479c-b434-737331f99d67 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.834782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.135153Z digest=sha256:ba60b647f62cfb07dd35bc66858187816f6d9db11388a9219017e60e1ece586c

Observation 292b5fe0-9da0-450f-a811-30b2b2ae16f9 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.815510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.256816Z digest=sha256:d2a908b498afd1ed70cdadd35c36e6577fb5e1d62b6cf2def861e821d6f03492

Observation 4399dc0f-e3cb-4c85-9f93-31825f001afe · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.798569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.327849Z digest=sha256:98120fa03886e9daae0c8077c47e7a4c552875535d03f588793936a8e21eac07

Observation 5fb6db72-521d-4eea-ba5d-8d15d5f7b6f5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.779688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.390660Z digest=sha256:656ed70d575ef78ad813c1876493f0fb665b68930292fc2da1f4aadc8caef7b6

Observation 16ced39f-51fd-4ba1-bc08-1f1ec6ad927e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.763152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.496477Z digest=sha256:e0edb370010fd0f7efad266edd3e9fa0c096c77b25cd76964f5c266898460983

Observation a5479ea5-f9f6-4c66-8327-1e0cbdf7ad6b · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.736635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.556587Z digest=sha256:bda0d4516f41ede5f8fb0664d90086b49cfccf45fa02fd76ac0c3bd50df1b1fe

Observation 334b0f3c-161e-4fa7-9fd9-51f853f72bb5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.712612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.621546Z digest=sha256:65aedad864adefb958ab84ebf9b53901de157bc2ca26ac84788c1d4ec6377bcc

Observation 2abaedfc-af62-469a-8c19-7697cfaef980 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.691201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.723436Z digest=sha256:c97d8455931f0cff5cacd04aee45194ca7ff50b69a01a77d7007008ca2d35a5f

Observation 8510c5d8-f4c5-495d-bbdf-be6ca6436edb · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.673643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.764626Z digest=sha256:8119cffa016cb89166890ad03736c844a99a3cd856a606fa58c4a3dd096a50f5

Observation fd9f8879-87e2-4a54-84b5-59f3da917bf5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.652297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.828019Z digest=sha256:3e2fbd80057f089c0e7489daaa50e69ee66ecad06180db786d61f1ebbd1ea0b8

Observation 22ced992-f3c1-4445-9f28-1447279b278e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.629924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.895073Z digest=sha256:f5e9bafcf372f7c9e834115abbee947121c3ea468893412252c33e5eb8bf9fba

Observation 9818ae67-df68-4042-a7d5-562a803eab9d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.600287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:57.960361Z digest=sha256:23d8662ae2f19b365693c993f74ee7826a97aca8149ed38c5ba1867fea1739cb

Observation ac6ce2f1-36ab-44f4-808a-f4d137033c05 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.578432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.024087Z digest=sha256:512be851571b10cdada8576ce01ac09348e326f67b0c2767e2ea744133837990

Observation d9984249-4631-4ce7-813a-f1ef90a0a32f · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.556838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.065219Z digest=sha256:bd5ca31aa9c87f0e1ae03d28dcd58902884c2f043bd9fad395504937102e5bd6

Observation 9a450b0f-a173-4834-a6fb-61b3e8b55c1d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.536723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.176510Z digest=sha256:f7767dbbb24a5ff7e845f96cc260e31c770db4061a180b37641bb4924c0fcdea

Observation dd41036f-1848-4dae-97e2-edb3d32b46d5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.521128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.244879Z digest=sha256:b864924a59812adba7559aa9fa84337d9a89e54cdfee66297660efd711cd4e74

Observation cc697d6f-033e-4ba5-b1f3-f299809ead8c · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.505202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.274042Z digest=sha256:21c183bc804ae1c810e30183dd2dec25a24476b15ae9e7c0c9644578c645c692

Observation c6a95066-2bf9-4323-96c3-827cc939b0a7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.487803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.279069Z digest=sha256:3c6b06ebc5befb6b003d6570fab6f96c57a68284add90d0887345aa0996843b4

Observation 7318b5a6-9db1-4156-a524-9e236f79ca00 · outbound

This paper cites We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.469236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.284468Z digest=sha256:6129efa5bb24d137202794929d9723de5e0344b0f59ec6bfc850b6f115b61e7c

Observation dd714861-6b4d-4d9c-b50a-12f83254acce · outbound

This paper cites 5th of October, there’s a significant event highlighted in blue - the launch.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 5th of October, there’s a significant event highlighted in blue - the launch

Reference 90

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:48:00.250320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.289503Z digest=sha256:8e7434b0eb6e49d1712eeaed45641c17189d7024fb23c7fe4a5a10d77b30facc

Observation f0a4420a-820d-4fcf-864c-64efaf7c91f2 · outbound

This paper cites Shared Prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shared Prompts

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.972954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.294865Z digest=sha256:ceca0665b54f7a80ec5d936a7084eed55d19ef75cdf4cb54cca889d9dad97295

Observation a91ca14b-6d26-4954-9d66-065ab404ec9f · outbound

This paper cites 8, the regional prompts obtain stronger responses in the corresponding local patches.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 8, the regional prompts obtain stronger responses in the corresponding local patches

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.679889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.298832Z digest=sha256:710567089693dce780f09ce6b1f8ed5a415a24702fd5f48cc69b91d977746ea1

Observation b917b074-82b1-4f5e-9acf-087729fe5ba7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:47:59.460818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.302953Z digest=sha256:5cf97bcc2d2feb367a54d7065cdea5d443d6d403a35fca1e767aba1e44984e28

Observation a663cbe5-11fc-4b8c-be9f-c887bd864bbe · outbound

This paper cites We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.238484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:47:58.308176Z digest=sha256:9864f21a55ed0a4ccb1d696a3c1aae523b36c24f5a1540f6ac747425524a7149

Pith citing papers

Observation 660bdf26-2245-4bdf-8e7e-b8d0ad692313 · inbound

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis cites this paper.

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:48:59.127431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T05:48:58.248509Z digest=sha256:55da73cd9ab99e4e188efcfa36b5e4f8fceb5946d8e7e3b3a6d98e09b84ca8f2