{"total":18,"items":[{"citing_arxiv_id":"2606.22481","ref_index":150,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Lighting-Consistent Object Transfer Across Radiance Fields","primary_cat":"cs.GR","submitted_at":"2026-06-21T12:50:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.18852","ref_index":45,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Aligning Implied Statements for Implicit Hate Speech Generalizability with Context-Bounded Semi-hard Negative Mining","primary_cat":"cs.CL","submitted_at":"2026-06-17T09:33:37+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"ImpSH improves cross-domain generalization in implicit hate speech classification by aligning posts with implied statements and applying context-bounded semi-hard negative mining within a triplet learning setup.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.17298","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins","primary_cat":"cs.CV","submitted_at":"2026-06-15T21:06:53+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"OR3 converts OR clips to action-driven digital twins, uses LLM imagination for hypothetical ActDTs, and achieves 57.6 R@1 and 77.3 R@5 on 276 implicit queries from 386 robotic knee procedure clips, outperforming baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.15604","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Parameter-Efficient Adaptation of SAM 3 for Automated ITV Generation from 4DCT Images","primary_cat":"eess.IV","submitted_at":"2026-06-14T05:12:57+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"LoRA-adapted SAM 3 with hard-negative mining and phase-coherent filtering achieves median Dice 0.968 on pulmonary structures from 4DCT using seven annotated volumes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02345","ref_index":36,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization","primary_cat":"stat.ML","submitted_at":"2026-06-01T14:54:29+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Sampling pairs directly with auxiliary information for higher inclusion probabilities on informative pairs yields near-full pairwise loss performance at reduced computational cost.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01482","ref_index":59,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG","primary_cat":"cs.CL","submitted_at":"2026-05-31T22:54:44+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"CERA fine-tunes a dense retriever with triplet contrastive learning plus attention alignment to human rationales, claiming better retrieval effectiveness and faithfulness on clinical trial reports than Contriever and standard hard-negative baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01334","ref_index":34,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"HOLA: Holistic Multi-Modal Alignment for Open-Set 3D Recognition","primary_cat":"cs.CV","submitted_at":"2026-05-31T16:36:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"HOLA introduces multi-view multi-text alignment and a decoupled contrastive loss for state-of-the-art open-vocabulary 3D recognition on long-tail benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01079","ref_index":18,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing","primary_cat":"cs.CV","submitted_at":"2026-05-31T07:54:26+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Chameleon proposes the first large-scale cross-domain compositing dataset and a disentangled encoder plus gated diffusion transformer that outperforms prior in-domain and cross-domain methods on plausibility and fidelity.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25598","ref_index":40,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SurfSurg6D: Geometry Consistent Dense Correspondence for Textureless Surgical Instrument Pose Estimation","primary_cat":"cs.CV","submitted_at":"2026-05-25T08:48:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A new synthetic dataset and geometry-consistent dense correspondence framework improve RGB-only pose estimation accuracy for surgical instruments on three evaluation datasets.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.23744","ref_index":36,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Contrast to Detect: Dynamic Graph Contrastive Regularization for Unsupervised Anomaly Detection in Multivariate Time Series","primary_cat":"cs.LG","submitted_at":"2026-05-22T15:18:53+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ContrastAD achieves highest mean F1 on all five MTS benchmarks and highest AUC on three by building DTW-based sparse graph snapshots and contrasting divergent pairs with a stable anchor instead of enforcing invariance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.19752","ref_index":53,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification","primary_cat":"cs.LG","submitted_at":"2026-05-19T12:19:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MSAlign aligns frozen DreaMS and ChemBERTa models with MLPs and candidate-based contrastive learning to outperform prior methods on molecule retrieval from MS/MS spectra while quantifying distribution shift in data splits.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.17854","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Learning over Positive and Negative Edges with Contrastive Message Passing","primary_cat":"cs.LG","submitted_at":"2026-05-18T04:52:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Contrastive Message Passing lets GNNs apply similarity-preserving transforms to positive edges and dissimilarity-inducing transforms to negative edges via soft positive semidefinite constraints on weights, yielding gains in low-label high-homophily regimes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10784","ref_index":52,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization","primary_cat":"cs.LG","submitted_at":"2026-05-11T16:18:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"MASS-DPO derives a Plackett-Luce-specific log-determinant Fisher information objective to select non-redundant negative samples, matching or exceeding multi-negative DPO performance with substantially fewer negatives across four benchmarks and three model families.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"DPO [ 42] aligns language models with human preferences58 by optimizing likelihood ratios of preferred over dispreferred responses, avoiding explicit reward59 modeling and associated complexities such as reward misgeneralization in RLHF [12, 38, 48]. Recent60 extensions include dynamic margins (ODPO; [3]) and prefix sharing for computational efficiency [52].61 However, standard DPO is restricted to binary preference pairs, limiting the diversity of supervision.62 Our approach extends beyond binary comparisons by leveraging actively selected, informative63 multi-negative samples.64 Multi-negative Preference Optimization.Recent work has extended standard DPO's binary65 preference pairs to leverage multiple negatives for richer comparative signals and enhanced alignment."},{"citing_arxiv_id":"2604.26057","ref_index":23,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection","primary_cat":"eess.AS","submitted_at":"2026-04-28T18:52:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Cosine similarity in SupCon with a delayed negative queue on wav2vec2 XLS-R yields the lowest equal error rates for deepfake audio detection on in-the-wild and pooled evaluations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.25273","ref_index":28,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval","primary_cat":"cs.CV","submitted_at":"2026-04-28T06:29:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SSA-ME uses saliency-aware modeling to reduce visual neglect and semantic drift, achieving SOTA results on the MMEB benchmark for multimodal retrieval.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.13313","ref_index":15,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding","primary_cat":"cs.LG","submitted_at":"2026-04-14T21:28:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Using lexical concreteness to guide contrastive negative mining and a new margin-based Cement loss, the Slipform framework reaches state-of-the-art on compositional benchmarks for vision-language models.","context_count":1,"top_context_role":"method","top_context_polarity":"use_method","context_text":"Despite these capabilities, VLMs are prone to biases stemming from limitations of the contrastive pretraining mechanism. These include failures to understand negations [9], object counting [10], spatial relations [11] and accurate entity associations [12]. One of the most significant limitations is bag-of-words-like behavior, which manifests as an inability to understand word order [13, 14] or attribute binding [15]. This limitation is primarily a byproduct of the contrastive objective. In the standard multimodal contrastive pretraining, training batches are curated by randomly sampled data points from the training dataset corpus. However, this random sampling often fails to provide effective negative samples to differentiate compositional semantics for each data point."},{"citing_arxiv_id":"2604.12443","ref_index":45,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DiffusionPrint: Learning Generative Fingerprints for Diffusion-Based Inpainting Localization","primary_cat":"cs.CV","submitted_at":"2026-04-14T08:30:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"DiffusionPrint learns robust forensic feature maps via MoCo-style contrastive training on diffusion inpainting fingerprints, boosting localization accuracy by up to 28% when fused into existing IFL systems and generalizing to unseen models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2411.18084","ref_index":27,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Exploration to Revelation: Detecting Dark Patterns in Mobile Apps","primary_cat":"cs.SE","submitted_at":"2024-11-27T06:39:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AppRay integrates LLM-guided task-oriented exploration with a contrastive learning multi-label classifier and rule-based refiner to detect intra- and inter-page dark patterns, reporting 0.89/0.85 F1 on new datasets with large gains over prior methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}