Pith. sign in

REVIEW 3 major objections 8 minor 1 cited by

Dynamic semantic channels remove loss discontinuities and raise tie-aware mAP for deep cross-modal hashing.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 11:31 UTC pith:PXDYEZ47

load-bearing objection Solid incremental fix to SCH’s discontinuous channels with honest multi-seed tables; the win-rate claim holds, the “continuity did it” story does not yet. the 3 major comments →

arxiv 2607.24567 v1 pith:PXDYEZ47 submitted 2026-07-27 cs.AI cs.CVcs.IR

DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

classification cs.AI cs.CVcs.IR
keywords deep semantic hashingcross-modal retrievalHamming spacesemantic channelstie-aware mAPloss landscapebinary codes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Deep semantic hashing turns images and text into short binary codes so that similar items sit close in Hamming space and can be retrieved quickly across modalities. Prior work set fixed-width target distance bands (semantic channels) from label similarity, but those bands jumped at zero similarity and put kinks in the loss surface. This paper replaces the fixed bands with continuously interpolated channel widths and positions that stay narrow for strong similarity and widen for weak similarity. Models trained with the new objective beat two strong baselines on most image–text retrieval tasks across two standard datasets, four code lengths, and two architectures. The authors also push for tie-aware mean average precision, because many database items share the same discrete Hamming distance from a query and ordinary mAP can change with arbitrary ordering inside those ties.

Core claim

A loss whose target Hamming channels have continuously varying width and left edge—both smooth functions of label cosine similarity—avoids the discontinuities of fixed-channel Semantic Channel Hashing and yields higher-quality binary codes. In 35 of 40 cross-modal and intra-modal retrieval tasks on MIR Flickr 25k and NUS-WIDE, at lengths 16–128 bits, models trained with this Dynamic Semantic Channel Hashing objective obtain higher tie-aware mAP than SCH and TDSRDH-loss, with gains up to 1.75 points over the second best.

What carries the argument

Dynamic Semantic Channel Hashing (DSCH) loss: pair-wise channel width w interpolates non-linearly between a minimum τ and the dissimilar-pair reserve via (1−S)^γ_w, the left edge p slides continuously with S, and the loss is the powered hinge outside that channel plus a light cross-modal quantization term.

Load-bearing premise

The measured gains come mainly from the continuous channel geometry itself, not from the accompanying changes in optimizer, batch size, image-augmentation schedule, or quantization weight.

What would settle it

Hold optimizer, batch size, augmentations, and quantization weight fixed and retrain the same architectures on the same splits with SCH versus DSCH; if DSCH no longer wins tie-aware mAP on a clear majority of the 40 tasks, the central claim is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The continuous-channel objective can be dropped into existing dual-stream and hybrid hashers without changing their architecture.
  • Tie-aware mAP should replace ordinary mAP for discrete Hamming retrieval so scores stop depending on arbitrary tie order.
  • Wider channels for weakly similar pairs give the model room for uncertain semantics instead of forcing a narrow band.
  • Larger hybrid backbones still improve under the same objective, indicating the loss is not tied to one model family within the tested range.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same continuous-channel construction could be tried in unsupervised hashing by replacing label cosine similarity with a soft affinity from a frozen teacher.
  • Separately ablating the progressive multi-image augmentations would show how much of the lift is geometric versus data-driven.
  • Distilling the larger hybrid models into the smaller dual-stream architecture under DSCH would test whether the loss preserves geometry under capacity reduction.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The manuscript proposes Dynamic Semantic Channel Hashing (DSCH), a supervised cross-modal hashing objective derived from Semantic Channel Hashing (SCH). Instead of SCH’s piecewise fixed-width target channels, DSCH assigns each pair a similarity-dependent channel width and left edge (Eqs. 6–8), intended to remove the discontinuity at zero label similarity. The objective is vectorized, combined with a cross-modal quantization term in Eq. (11), and evaluated with a tie-aware mAP metric adapted from prior work. Experiments cover MIR Flickr 25k and NUS-WIDE, TDSRDH and a CLIP-based architecture, code lengths 16–128, and at least three runs per setting. The authors report DSCH as best in 35 of 40 retrieval settings, with gains up to 1.75 percentage points.

Significance. If the comparative result holds, DSCH is a useful, simple modification of semantic-channel hashing, and the tie-aware mAP discussion addresses a real evaluation problem caused by tied Hamming distances. Notable strengths are the explicit iterative and vectorized formulations, a public reference implementation of the losses and metrics, a largely shared optimizer/batch/augmentation recipe across compared losses, multi-seed reporting, and evaluation on two datasets, two architectures, and four code lengths. The validation-query/validation-retrieval split is also a sound addition. The present evidence is stronger for “DSCH performs well under this common recipe” than for the stronger conclusion that continuity removal is the causal source of the gains.

major comments (3)
  1. [Abstract; Tables VI and IX] The abstract claims “significantly higher” scores in 35/40 tasks, but no significance test is reported. At the precision shown, MIR Flickr T→T at k=32 is a tie (DSCH 70.80±0.25 vs SCH 70.80±0.35), so the tables show 34 strict wins, one tie, and five losses, not 35 strictly higher scores. Several margins also overlap the reported run deviations, e.g. MIR Flickr I→T k=16 (84.19±0.37 vs 83.73±0.83). Please define the ± quantity and number of seeds, report paired tests or confidence intervals/effect sizes across matched runs, address multiple comparisons, and revise the win count and “significantly” wording accordingly.
  2. [§III-C, Eqs. (5)–(11); §IV-C; Tables VI and IX] The manuscript attributes the gains to removing the S_ij=0 discontinuity via Eqs. (6)–(8), but the evaluated DSCH system in Eq. (11) changes several components simultaneously: dynamic channels, the added cross-modal quantization term Lq with κq=0.01, and τ=1 rather than SCH’s τ=3. The SCH comparison in §IV-C appears to use Eq. (5) without Lq. Since Eq. (10) explicitly aligns both modalities with a shared quantized target, it may account for part of the cross-modal gains. A component ablation—DSCH without Lq, SCH plus the same Lq, DSCH with τ=3, and ideally dynamic width and dynamic position separated—is needed to support the mechanistic explanation.
  3. [§V, Tables VII–VIII; §IV-C] The validation mAP results for γw∈{6,8,12,14} are mostly within the reported deviations, and γw=8 is then selected using ROC-AUC, a different criterion from the paper’s primary metric. Even Table VIII is not uniformly best at 8 (e.g. MIR Flickr I→I favors 12). Please state an a priori aggregation/selection criterion, report mAP sensitivity at the final test setting, or use nested validation. Relatedly, SCH retains τ=3 and the reimplemented TDSRDH-loss uses altered weights αTDSRDH=0.05, βTDSRDH=1; comparable validation tuning or a clear justification for this asymmetry is needed for a fair loss comparison.
minor comments (8)
  1. [§III-A] l_i∈{0,1}^c is repeatedly called “one-hot,” but the datasets are multi-label. “Binary indicator” or “multi-hot” vector would be accurate.
  2. [Eq. (9)] “Haramad product” should be “Hadamard product.”
  3. [Table I; Eq. (11)] n is defined in Table I as |D|, but Eq. (11) uses n for batch-size normalization. Please introduce a distinct batch-size symbol and state whether ordered pairs, i=j terms, and all four modality combinations are included.
  4. [Eq. (15)] The trailing “o/w,” in the piecewise definition is undefined and should be replaced by “otherwise” or removed if unintended.
  5. [§IV-B] The claim that ref. [15] “cannot be used in a cross-modal retrieval setting” appears inconsistent with that paper’s stated cross-modal hashing task. Please clarify the intended architectural distinction and temper the novelty statement for CLIPHash.
  6. [Figure 1] Figure 1 compares SCH at τ=3 with DSCH at γw=2, τ=1, whereas the main experiments use γw=8. Please mark the figure as illustrative and include the experimental parameterization or a second panel.
  7. [Tables VI and IX] Table IX repeats the TDSRDH k=32 results from Table VI. Please explicitly state that the 40-task count consists of the 32 Table VI settings plus eight new CLIPHash settings, so the totals are easy to audit.
  8. [§VI] Typographical issue: “dynamic target sematic channel width” should be “semantic.”

Circularity Check

0 steps flagged

No circularity: DSCH is a designed empirical loss objective, not a derivation that rewrites fitted inputs as predictions.

full rationale

This is a standard empirical deep-learning methods paper. The central object (DSCH, Eqs. 6–11) is an explicitly designed training loss that interpolates semantic-channel width and position from label cosine similarity S_ij; it does not claim to derive or predict a physical/mathematical quantity from first principles. Reported gains are experimental comparisons of tie-aware mAP (taken from He et al. / McSherry & Najork) under shared training recipes against SCH and TDSRDH-loss on two public datasets and two architectures (Tables VI, IX). Hyperparameters (γ_w, τ, κ_q) are set by hand or chosen on a validation sweep—normal free-parameter practice, not a fitted quantity renamed as a test prediction. The SCH foundation is cited from external authors (Hu et al.), not a self-citation uniqueness theorem. No equation forces the reported mAP by construction; no load-bearing self-citation chain; no renaming of a known closed-form result. Attribution concerns (missing ablations of L_q vs. continuity, post-hoc γ_w via ROC-AUC) are soundness issues, not circularity. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 1 invented entities

The claim rests on standard supervised multi-label similarity, the SCH geometric picture of target Hamming bands, and a handful of hand-chosen or validation-chosen scalars that shape channel width and quantization strength. No new physical entities; ‘dynamic semantic channels’ are a loss-design construct built on Hu et al.

free parameters (6)
  • γ_w (channel width curve modifier) = 8
    Controls non-linear width interpolation; selected as 8 after validation sweep where mAP was inconclusive and ROC-AUC preferred 8.
  • τ (minimum channel width) = 1
    Floor on semantic channel width; set to 1 vs SCH’s fixed width 3 without exhaustive justification beyond continuity and uncertainty narrative.
  • λ_neg (minimum distance for negatives) = k/2
    Fixed to k/2 following SCH; shapes the entire channel layout.
  • κ_q (quantization loss weight) = 0.01
    Balances Lq against LD; set to 0.01 by stability/performance preference.
  • γ_ℓ, α, β (loss curve and set weights) = 1, 1, 1
    Defaulted to 1 (linear, unweighted); available as knobs but not the main reported setting.
  • TDSRDH baseline loss weights α_TDSRDH, β_TDSRDH = 0.05, 1
    Retuned to 0.05 and 1 under Adam/normalization so the re-implementation matches claimed performance—affects fairness of comparison.
axioms (5)
  • domain assumption Two samples are treated as similar for evaluation iff they share at least one label (δ = 1[S_ij > 0]); graded cosine similarity of multi-hot labels is the right continuous similarity for positioning channels.
    Stated in §III-A and used in Eqs. 1, 6–8 and in mAP definitions; standard in this literature but not validated against other similarity graphs.
  • domain assumption Hamming distance may be computed from cosine similarity of real-valued pre-quantization codes during training (Eq. 2).
    Required for differentiability; inherited from SCH-style deep hashing practice (§III-A/B).
  • ad hoc to paper A continuous loss landscape without the S_ij=0 jump is preferable for optimization and motivates dynamic width/position.
    Design goal in §III-C; supported by landscape plots but not by a controlled ablation that isolates discontinuity removal from other recipe changes.
  • domain assumption Tie-aware AP/mAP is the appropriate primary metric because discrete Hamming distances induce ties.
    §IV-D citing He et al. and McSherry & Najork; reasonable and improves comparability.
  • domain assumption Random fixed-seed subsampling of query/train sets (without iterative stratification) adequately represents the multi-label distribution.
    §IV-A; authors argue against per-label balanced sampling but this choice affects all reported numbers.
invented entities (1)
  • Dynamic semantic channel (width w_ij and left edge p_ij) no independent evidence
    purpose: Replace SCH’s fixed-width, discontinuous target bands with a continuous, similarity-dependent target region in Hamming space.
    Methodological construct defined in Eqs. 6–7; builds directly on Hu et al.’s semantic channels rather than introducing a new scientific object with external ontology.

pith-pipeline@v1.2.0-grok45-kimik3 · 25942 in / 3713 out tokens · 79496 ms · 2026-07-31T11:31:25.624424+00:00 · methodology

0 comments
read the original abstract

Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and introduced a loss function based on predefined so-called semantic channels with fixed width and Hamming distances derived from label similarities. However, this formulation also introduced discontinuities into the loss landscape, complicating optimization. Based on these observations, we propose a newly designed loss function, Dynamic Semantic Channel Hashing (DSCH), using dynamically sized and positioned semantic channels in order to avoid loss landscape discontinuities. Furthermore, we endorse the use of tie-aware Mean Average Precision (mAP) as evaluation metric as it addresses the ambiguity in sample retrieval ordering, which emerges from the discreteness of hash code distances. Finally, multiple experimental settings conducted on two popular datasets and incorporating two different model architectures provide strong evidence that training using the DSCH objective outperforms training using other state-of-the-art loss functions. In a total of 35 out of 40 cross-modal and intra-modal retrieval tasks, models trained with DSCH achieve significantly higher tie-aware mAP scores across all four tested hash code lengths, showing compelling results across model architecture and used dataset. The mAP score uplifts are consistent and amount up to 1.75 percentage points compared to the respective second best.

Figures

Figures reproduced from arXiv: 2607.24567 by Christian Bergler, Christian Riess, Daniel Loebenberger, Tobias J. Bauer.

Figure 1
Figure 1. Figure 1: Visual comparison of the loss landscapes of Semantic Channel Hashing [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Schematic representation of the dataset partitioning process to obtain [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Schematic overview over the two model architectures employed for the experiments in this paper. Shaded in purple are off-the-shelf pre-trained [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Demonstration of different retrieval sequences that are semantically [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CosmoLattice 2.0

    astro-ph.CO 2026-07 accept novelty 6.0

    CosmoLattice v2.0 extends lattice cosmology simulations with non-minimal scalars, ALP–gauge couplings, defect networks, low-storage RK integrators, optimized GWs, and O(10) GPU speedups.

Reference graph

Works this paper leans on

42 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    A Survey on Deep Hashing Methods,

    X. Luo et al., “A Survey on Deep Hashing Methods,”ACM Transac- tions on Knowledge Discovery from Data, vol. 17, no. 1, pp. 1–50, Feb. 28, 2023.DOI: 10.1145/3532624

  2. [2]

    Learning to hash: A comprehensive survey of deep learning-based hashing methods,

    A. Singh and S. Gupta, “Learning to hash: A comprehensive survey of deep learning-based hashing methods,”Knowledge and Information Systems, vol. 64, no. 10, pp. 2565–2597, Oct. 2022.DOI: 10.1007/ s10115-022-01734-0

  3. [3]

    Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions,

    T. Wang, F. Li, L. Zhu, J. Li, Z. Zhang, and H. T. Shen, “Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions,” Proceedings of the IEEE, vol. 112, no. 11, pp. 1716–1754, Nov. 2024. DOI: 10.1109/JPROC.2024.3525147

  4. [4]

    The State of the Art for Cross-Modal Retrieval: A Survey,

    K. Zhou, F. H. Hassan, and G. K. Hoon, “The State of the Art for Cross-Modal Retrieval: A Survey,”IEEE Access, vol. 11, pp. 138 568– 138 589, 2023.DOI: 10.1109/ACCESS.2023.3338548

  5. [5]

    Deep Hashing for Scalable Image Search,

    J. Lu, V . E. Liong, and J. Zhou, “Deep Hashing for Scalable Image Search,”IEEE Transactions on Image Processing, vol. 26, no. 5, pp. 2352–2367, May 2017.DOI: 10.1109/TIP.2017.2678163

  6. [6]

    Robust and Secure Image Fingerprinting Learned by Neural Network,

    Y . Li, D. Wang, and L. Tang, “Robust and Secure Image Fingerprinting Learned by Neural Network,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 2, pp. 362–375, Feb. 2020. DOI: 10.1109/TCSVT.2019.2890966

  7. [7]

    Deep Semantic Reconstruction Hashing for Similarity Retrieval,

    Y . Wang, X. Ou, J. Liang, and Z. Sun, “Deep Semantic Reconstruction Hashing for Similarity Retrieval,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 1, pp. 387–400, Jan. 2021. DOI: 10.1109/TCSVT.2020.2974768

  8. [8]

    TransHash: Transformer-based Hamming Hashing for Efficient Image Retrieval,

    Y . Chen, S. Zhang, F. Liu, Z. Chang, M. Ye, and Z. Qi, “TransHash: Transformer-based Hamming Hashing for Efficient Image Retrieval,” inProceedings of the 2022 International Conference on Multimedia Retrieval, Newark NJ USA: ACM, Jun. 27, 2022, pp. 127–136.DOI: 10.1145/3512527.3531405

  9. [9]

    Deep Semantic Hashing Using Pairwise Labels,

    R. Xuan, J. Shim, and S. -G. Lee, “Deep Semantic Hashing Using Pairwise Labels,”IEEE Access, vol. 9, pp. 91 934–91 949, 2021.DOI: 10.1109/ACCESS.2021.3092150

  10. [10]

    Deep Cross-Modal Hashing,

    Q.-Y . Jiang and W.-J. Li, “Deep Cross-Modal Hashing,” in2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI: IEEE, Jul. 2017, pp. 3270–3278.DOI: 10.1109/CVPR. 2017.348

  11. [11]

    Deep Cross-modal Hashing Retrieval Based on Semantics Preserving and Vision Transformer,

    J. Hong and H. Liu, “Deep Cross-modal Hashing Retrieval Based on Semantics Preserving and Vision Transformer,” inProceedings of the 2022 6th International Conference on Electronic Information Technology and Computer Engineering, Xiamen China: ACM, Oct. 21, 2022, pp. 52–57.DOI: 10.1145/3573428.3573439

  12. [12]

    TECMH: Transformer- Based Cross-Modal Hashing For Fine-Grained Image-Text Retrieval,

    Q. Li, L. Ma, Z. Jiang, M. Li, and B. Jin, “TECMH: Transformer- Based Cross-Modal Hashing For Fine-Grained Image-Text Retrieval,” Computers, Materials & Continua, vol. 75, no. 2, pp. 3713–3728, 2023.DOI: 10.32604/cmc.2023.037463

  13. [13]

    Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals,

    L. Jin, Z. Li, and J. Tang, “Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals,”IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 4, pp. 1838–1851, Apr. 2023.DOI: 10.1109/TNNLS.2020.2997020

  14. [14]

    Transformer-Based Discriminative and Strong Rep- resentation Deep Hashing for Cross-Modal Retrieval,

    S. Zhou et al., “Transformer-Based Discriminative and Strong Rep- resentation Deep Hashing for Cross-Modal Retrieval,”IEEE Access, vol. 11, pp. 140 041–140 055, 2023.DOI: 10.1109/ACCESS.2023. 3339581

  15. [15]

    When CLIP meets cross- modal hashing retrieval: A new strong baseline,

    X. Xia, G. Dong, F. Li, L. Zhu, and X. Ying, “When CLIP meets cross- modal hashing retrieval: A new strong baseline,”Information Fusion, vol. 100, p. 101 968, Dec. 2023.DOI: 10.1016/j.inffus.2023.101968

  16. [16]

    Similarity Preserving Transformer Cross-Modal Hashing for Video-Text Re- trieval,

    Q. Huang, S. Peng, X. Shen, Y . -H. Yuan, and S. Pan, “Similarity Preserving Transformer Cross-Modal Hashing for Video-Text Re- trieval,” inProceedings of the 32nd ACM International Conference on Multimedia, Melbourne VIC Australia: ACM, Oct. 28, 2024, pp. 5883–5891.DOI: 10.1145/3664647.3681606

  17. [17]

    CKDH: CLIP- Based Knowledge Distillation Hashing for Cross-Modal Retrieval,

    J. Li, W. K. Wong, L. Jiang, X. Fang, S. Xie, and Y . Xu, “CKDH: CLIP- Based Knowledge Distillation Hashing for Cross-Modal Retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6530–6541, Jul. 2024.DOI: 10.1109/TCSVT.2024. 3350695

  18. [18]

    Cross-Modal Hashing Method With Properties of Hamming Space: A New Perspective,

    Z. Hu, Y .-M. Cheung, M. Li, and W. Lan, “Cross-Modal Hashing Method With Properties of Hamming Space: A New Perspective,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 7636–7650, Dec. 2024.DOI: 10.1109/TPAMI.2024.3392763

  19. [19]

    CLIP Multi-modal Hashing for Multimedia Retrieval

    J. Zhu et al. “CLIP Multi-modal Hashing for Multimedia Retrieval.” arXiv: 2410.07783[cs]

  20. [20]

    Convolutional networks and applications in vision,

    Y . LeCun, K. Kavukcuoglu, and C. Farabet, “Convolutional networks and applications in vision,” inProceedings of 2010 IEEE International Symposium on Circuits and Systems, Paris, France: IEEE, May 2010, pp. 253–256.DOI: 10.1109/ISCAS.2010.5537907

  21. [21]

    ImageNet Classifi- cation with Deep Convolutional Neural Networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classifi- cation with Deep Convolutional Neural Networks,” inAdvances in Neural Information Processing Systems, vol. 25, Curran Associates, Inc., 2012

  22. [22]

    Deep Residual Learning for Image Recognition

    K. He, X. Zhang, S. Ren, and J. Sun. “Deep Residual Learning for Image Recognition.” arXiv: 1512.03385[cs]

  23. [23]

    Triplet-Based Deep Hashing Network for Cross-Modal Retrieval,

    C. Deng, Z. Chen, X. Liu, X. Gao, and D. Tao, “Triplet-Based Deep Hashing Network for Cross-Modal Retrieval,”IEEE Transactions on Image Processing, vol. 27, no. 8, pp. 3893–3903, Aug. 2018.DOI: 10.1109/TIP.2018.2821921

  24. [24]

    CLIP-based fusion-modal reconstructing hashing for large-scale unsupervised cross- modal retrieval,

    L. Mingyong, L. Yewen, G. Mingyuan, and M. Longfei, “CLIP-based fusion-modal reconstructing hashing for large-scale unsupervised cross- modal retrieval,”International Journal of Multimedia Information Retrieval, vol. 12, no. 1, p. 2, Jun. 2023.DOI: 10.1007/s13735-023- 00268-7

  25. [25]

    Attention Is All You Need

    A. Vaswani et al. “Attention Is All You Need.” arXiv: 1706.03762 [cs]

  26. [26]

    Advancements in natural language processing: Implications, challenges, and future directions,

    Supriyono, A. P. Wibawa, Suyono, and F. Kurniawan, “Advancements in natural language processing: Implications, challenges, and future directions,”Telematics and Informatics Reports, vol. 16, p. 100 173, Dec. 2024.DOI: 10.1016/j.teler.2024.100173

  27. [27]

    Hashing as Tie-Aware Learning to Rank

    K. He, F. Cakir, S. A. Bargal, and S. Sclaroff. “Hashing as Tie-Aware Learning to Rank.” arXiv: 1705.08562[stat]

  28. [28]

    NUS- WIDE: A real-world web image database from National University 12 of Singapore,

    T.-S. Chua, J. Tang, R. Hong, H. Li, Z. Luo, and Y . Zheng, “NUS- WIDE: A real-world web image database from National University 12 of Singapore,” inProceedings of the ACM International Conference on Image and Video Retrieval, Santorini, Fira Greece: ACM, Jul. 8, 2009, pp. 1–9.DOI: 10.1145/1646396.1646452

  29. [29]

    The MIR flickr retrieval evaluation,

    M. J. Huiskes and M. S. Lew, “The MIR flickr retrieval evaluation,” in Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval, Vancouver British Columbia Canada: ACM, Oct. 30, 2008, pp. 39–43.DOI: 10.1145/1460096.1460104

  30. [30]

    On the Stratification of Multi-label Data,

    K. Sechidis, G. Tsoumakas, and I. Vlahavas, “On the Stratification of Multi-label Data,” inMachine Learning and Knowledge Discovery in Databases, D. Gunopulos, T. Hofmann, D. Malerba, and M. Vazirgiannis, Eds., vol. 6913, Berlin, Heidelberg: Springer Berlin Heidelberg, 2011, pp. 145–158.DOI: 10.1007/978-3-642-23808-6 10

  31. [31]

    Efficient Estimation of Word Representations in Vector Space

    T. Mikolov, K. Chen, G. Corrado, and J. Dean. “Efficient Estimation of Word Representations in Vector Space.” version 3

  32. [32]

    Garbe,Symspellpy: Python SymSpell, version 6.9.0, Mar

    mammothb and W. Garbe,Symspellpy: Python SymSpell, version 6.9.0, Mar. 9, 2025

  33. [33]

    Learning Transferable Visual Models From Natural Language Supervision,

    A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,”Proceedings of the 38th International Conference on Machine Learning, vol. 139, pp. 8748–8763, Jul. 2021

  34. [34]

    Data Filtering Networks

    A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. Toshev, and V . Shankar. “Data Filtering Networks.” arXiv: 2309.17425[cs]

  35. [35]

    Ilharco et al.,OpenCLIP, version 3.1, Zenodo, Jul

    G. Ilharco et al.,OpenCLIP, version 3.1, Zenodo, Jul. 28, 2021.DOI: 10.5281/ZENODO.5143773

  36. [36]

    CLIP4Hashing: Unsupervised Deep Hashing for Cross-Modal Video-Text Retrieval,

    Y . Zhuo, Y . Li, J. Hsiao, C. Ho, and B. Li, “CLIP4Hashing: Unsupervised Deep Hashing for Cross-Modal Video-Text Retrieval,” inProceedings of the 2022 International Conference on Multimedia Retrieval, ser. ICMR ’22, New York, NY , USA: Association for Computing Machinery, Jun. 27, 2022, pp. 158–166.DOI: 10.1145/ 3512527.3531381

  37. [37]

    CCAH: A CLIP-Based Cycle Align- ment Hashing Method for Unsupervised Vision-Text Retrieval,

    M. Li, L. Ma, Y . Li, and M. Ge, “CCAH: A CLIP-Based Cycle Align- ment Hashing Method for Unsupervised Vision-Text Retrieval,”Inter- national Journal of Intelligent Systems, vol. 2023, no. 1, p. 7 992 047, 2023.DOI: 10.1155/2023/7992047

  38. [38]

    A Highly Efficient Zero- Shot Cross-Modal Hashing Method Based on CLIP,

    L. Cao, H. Xiao, W. Song, and H. Li, “A Highly Efficient Zero- Shot Cross-Modal Hashing Method Based on CLIP,” in2024 5th International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT), Mar. 2024, pp. 868–873.DOI: 10. 1109/AINIT61980.2024.10581810

  39. [39]

    Adam: A Method for Stochastic Optimiza- tion

    D. P. Kingma and J. Ba. “Adam: A Method for Stochastic Optimiza- tion.” arXiv: 1412.6980[cs]

  40. [40]

    Deep Semantic Hashing with Generative Adversarial Networks

    Z. Qiu, Y . Pan, T. Yao, and T. Mei. “Deep Semantic Hashing with Generative Adversarial Networks.” arXiv: 1804.08275[cs]

  41. [41]

    A Comprehensive Survey of Image Augmentation Techniques for Deep Learning,

    M. Xu, S. Yoon, A. Fuentes, and D. S. Park, “A Comprehensive Survey of Image Augmentation Techniques for Deep Learning,”Pattern Recognition, vol. 137, p. 109 347, May 2023.DOI: 10.1016/j.patcog. 2023.109347

  42. [42]

    Computing Information Retrieval Performance Measures Efficiently in the Presence of Tied Scores,

    F. McSherry and M. Najork, “Computing Information Retrieval Performance Measures Efficiently in the Presence of Tied Scores,” inAdvances in Information Retrieval, C. Macdonald, I. Ounis, V . Plachouras, I. Ruthven, and R. W. White, Eds., vol. 4956, Berlin, Heidelberg: Springer Berlin Heidelberg, 2008, pp. 414–421.DOI: 10.1007/978-3-540-78646-7 38