Replacing an object's name in a caption with the image tokens of that object during pretraining gives explicit object-entity grounding, making MLLM alignment several times more data-efficient and boosting grounding and perception scores.
hub
Title resolution pending
7 Pith papers cite this work, alongside 1,978 external citations. Polarity classification is still indexing.
hub tools
verdicts
CONDITIONAL 7representative citing papers
PipeMFL-240K provides 249,320 labeled MFL pipeline scans across 12 categories — the first large public benchmark for MFL defect detection — and baseline tests show current detectors cap near 0.5 mAP50.
VolSegGS reconstructs dynamic volumetric scenes from rendered images with deformable 3D Gaussians and enables real-time interactive segmentation and tracking of regions over time.
DRAFE, an asymmetric fusion ensemble of two LW-DETR and one RF-DETR detectors, achieved 0.4022 mAP and sixth place on AI City Challenge 2026 Track 6 via class-consistent matching and complementary recovery.
A wavelet-guided neural pipeline recovers previously known narrowband radio events from FAST observations of 33 exoplanet systems and reduces 139,127 detections to 803 veto-ready candidates; its one new candidate, toward K2-155, is argued to be terrestrial RFI.
Frozen 1120x1120 inference of a 704-trained RF-DETR achieved +0.0382 AP over the 704 baseline on an aggregate-only hidden cross-city benchmark, and a fine-tuning run improved in-domain validation AP while its hidden AP did not rise.
On Moon and Mars imagery, YOLO yields the most balanced crater detection, while ResNet-50 excels at large craters, under a two-stage CNN-based framework.
citing papers explorer
-
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
Replacing an object's name in a caption with the image tokens of that object during pretraining gives explicit object-entity grounding, making MLLM alignment several times more data-efficient and boosting grounding and perception scores.
-
PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging
PipeMFL-240K provides 249,320 labeled MFL pipeline scans across 12 categories — the first large public benchmark for MFL defect detection — and baseline tests show current detectors cap near 0.5 mAP50.
-
VolSegGS: Segmentation and Tracking in Dynamic Volumetric Scenes via Deformable 3D Gaussians
VolSegGS reconstructs dynamic volumetric scenes from rendered images with deformable 3D Gaussians and enables real-time interactive segmentation and tracking of regions over time.
-
DRAFE: Domain-Robust Asymmetric Fusion of Heterogeneous Detection Transformers for Cross-City Fine-Grained Traffic Object Detection
DRAFE, an asymmetric fusion ensemble of two LW-DETR and one RF-DETR detectors, achieved 0.4022 mAP and sixth place on AI City Challenge 2026 Track 6 via class-consistent matching and complementary recovery.
-
A Wavelet-Integrated Search Pipeline for Narrowband Technosignatures in FAST Observations of 33 Exoplanet Systems
A wavelet-guided neural pipeline recovers previously known narrowband radio events from FAST observations of 33 exoplanet systems and reduces 139,127 detections to 803 veto-ready candidates; its one new candidate, toward K2-155, is argued to be terrestrial RFI.
-
Frozen High-Resolution Inference for Cross-City Object Detection: An AI City Challenge 2026 Study
Frozen 1120x1120 inference of a 704-trained RF-DETR achieved +0.0382 AP over the 704 baseline on an aggregate-only hidden cross-city benchmark, and a fine-tuning run improved in-domain validation AP while its hidden AP did not rise.
-
Deep learning framework for crater detection and identification on the Moon and Mars
On Moon and Mars imagery, YOLO yields the most balanced crater detection, while ResNet-50 excels at large craters, under a two-stage CNN-based framework.