REVIEW 4 major objections 63 references
MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering
T0 review · 4 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read MobQA, a 5,800-pair benchmark of GPS-based question answering, shows large language models can retrieve mobility facts but struggle to interpret what the movement means.
desk verdict MobQA abstract looks sensible, but the full text is an unrelated remote-sensing paper, so the submission is incoherent and unverdictable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
MobQA itself — a dataset of 5,800 question-answer pairs over daily-to-weekly GPS trajectories, organized into factual retrieval, multiple-choice reasoning, and free-form explanation items that jointly require spatial, temporal, and semantic reasoning. The benchmark is the instrument that carries the claim: it separates retrieval from inference and interpretation so that LLM performance differences across question types read as differences in semantic understanding.
What would settle it
Take a random subset of the 5,800 pairs, have independent annotators re-create gold answers and score the same model outputs by human judgment; if human scores disagree with the paper's reported ordering across question types, or if inter-annotator agreement on the gold answers is low, the central claim collapses.
Extended reading notes
Core claim
The paper's central claim is that semantic understanding of mobility data can be decomposed into three layers and tested by question answering: factual retrieval (what happened), multiple-choice reasoning (why it happened), and free-form explanation (what the pattern means). On MobQA, major LLMs show strong performance on the first layer but significant limitations on the second and third, and trajectory length substantially affects effectiveness. The authors present this as evidence that state-of-the-art LLMs have not yet achieved robust semantic mobility understanding.
Load-bearing premise
The central finding stands only if the 5,800 question-answer pairs have correct, unbiased gold answers scored by a reliable, non-contaminated protocol — and the supplied full text is a different paper, so none of these controls can be verified from this submission.
Editorial extensions
If this is right
- LLMs already handle factual extraction from mobility traces well, so future work should shift toward reasoning and interpretation rather than basic retrieval.
- Semantic reasoning and free-form explanation are the current bottlenecks in LLM mobility understanding.
- Trajectory length matters: longer trajectories degrade model effectiveness, so compressing or hierarchically summarizing long GPS traces could improve performance.
- MobQA provides a reusable evaluation suite for any future model that claims semantic understanding of human mobility data.
Reading between the lines
- If the reported failure is driven by trajectory length, then retrieval-augmented or hierarchical summarization of GPS traces might improve LLM reasoning — a testable extension the paper itself does not run.
- Human performance on the same 5,800 pairs would calibrate what 'semantic understanding' should mean for mobility and whether the gap the paper reports is a model limitation or a task difficulty artifact.
- The multiple-choice and free-form layers of MobQA could double as a general probe of spatial-temporal reasoning in LLMs, not just mobility-specific understanding.
- Because the explanation questions are free-form, the choice of automatic scoring metric likely changes the leaderboard; a validated human-eval subset would be needed before using MobQA as a competitive ranking instrument.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, as reviewed, consists of an abstract announcing MobQA, a benchmark of 5,800 question-answer pairs over human GPS trajectories, with three question types (factual retrieval, multiple-choice reasoning, free-form explanation) and an evaluation of major LLMs reporting strong factual retrieval but weak reasoning/explanation and a trajectory-length effect. The supplied full text, however, is an entirely different paper: "VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection" (VG-DETR, arXiv:2508.11167v2). The full text contains no mention of MobQA, human mobility, GPS trajectories, question-answer construction, LLM evaluation, or free-form answer scoring. Consequently, every empirical claim in the abstract is unsupported by the reviewed manuscript.
Significance. A well-constructed mobility QA benchmark would be a useful contribution, and the abstract gestures at a concrete resource (a GitHub URL, 5,800 pairs, three question types). However, significance can only be assessed if the benchmark construction and evaluation are described and checkable. The current submission provides no such description, no protocol, no annotation-quality evidence, and no experimental details. The full text is a remote sensing object detection paper, so the benchmark and its findings cannot be evaluated. The paper's potential value is real but entirely unverified in the submitted material.
major comments (4)
- [Abstract vs. Full Text] The abstract describes MobQA, LLM evaluation, and trajectory-length effects, but the full text is VG-DETR, a source-free remote sensing object detection paper. The running header identifies arXiv:2508.11167v2, and the body discusses pseudo-labels, DETR, and domain adaptation. There is no overlap of topic. The central claim of the abstract is therefore unsupported by the reviewed text.
- [No benchmark construction details] The abstract states that MobQA comprises 5,800 high-quality question-answer pairs, but the manuscript gives no information about how trajectories were collected, how questions were generated, how gold answers were verified, whether expert review or inter-annotator agreement was used, or what quality controls produced the asserted 'high-quality' label. Without this, the benchmark's trustworthiness is unassessable.
- [No evaluation protocol] The abstract claims strong factual retrieval but weak reasoning and explanation performance across 'major LLMs.' The full text provides no model list, prompt templates, answer-scoring method (especially for free-form explanations), variance or error bars, or human agreement data. It is therefore impossible to verify the direction or magnitude of the claimed effects, or whether an LLM-based judge might be measuring self-agreement rather than semantic understanding.
- [No contamination or representativeness controls] The claim that MobQA measures general LLM semantic understanding of mobility data presupposes that the test items are not memorized from pretraining and that trajectories and questions are representative. The manuscript provides no discussion of contamination checks, no diversity analysis of trajectories, and no evidence that the benchmark supports the broad conclusions stated in the abstract.
Circularity Check
No circular derivation is exhibitable from the MobQA abstract; the supplied full text is an unrelated paper, making the benchmark claims unverifiable but not circular.
full rationale
The claimed contribution of the MobQA abstract is an externally constructed benchmark (5,800 human-authored question-answer pairs over GPS trajectories) plus an evaluation of existing LLMs against that benchmark. There is no derivation chain in which an output is defined in terms of an input, a fitted parameter is renamed as a prediction, or a self-citation is load-bearing. The abstract does not specify how free-form explanations were scored, so one cannot quote any statement that they were graded by the same LLMs under evaluation; flagging that as circular would be speculation. The full text supplied with the submission is a different paper (VG-DETR, arXiv:2508.11167v2) with no mention of MobQA, trajectories, or QA construction. That mismatch prevents checking the benchmark's quality controls, gold-answer consistency, contamination safeguards, or scoring protocol, but missing methodology is a completeness/integrity concern, not a circularity reduction. Under the hard rule that circularity may only be claimed with a quoted reduction or fitted-parameter-as-prediction, no circular step can be identified. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption QA accuracy on MobQA is a valid operationalization of semantic understanding of mobility data
- domain assumption The 5,800 question-answer pairs have correct gold answers and were quality-controlled
- domain assumption Free-form explanation answers are scored reliably and fairly
Cite this review
Pith. "Pith review of MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering." pith.science (2026). https://pith.science/paper/C7E6MFI2
@misc{pith2026250811163,
author = {Pith},
title = {Pith review of: MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/C7E6MFI2}},
note = {Machine review of arXiv:2508.11163}
}
read the original abstract
This paper presents MobQA, a benchmark dataset designed to evaluate the semantic understanding capabilities of large language models (LLMs) for human mobility data through natural language question answering. While existing models excel at predicting human movement patterns, it remains unobvious how much they can interpret the underlying reasons or semantic meaning of those patterns. MobQA provides a comprehensive evaluation framework for LLMs to answer questions about diverse human GPS trajectories spanning daily to weekly granularities. It comprises 5,800 high-quality question-answer pairs across three complementary question types: factual retrieval (precise data extraction), multiple-choice reasoning (semantic inference), and free-form explanation (interpretive description), which all require spatial, temporal, and semantic reasoning. Our evaluation of major LLMs reveals strong performance on factual retrieval but significant limitations in semantic reasoning and explanation question answering, with trajectory length substantially impacting model effectiveness. These findings demonstrate the achievements and limitations of state-of-the-art LLMs for semantic mobility understanding.\footnote{MobQA dataset is available at https://github.com/CyberAgentAILab/mobqa.}
Reference graph
Works this paper leans on
-
[1]
Instance Feature Visualization: We evaluate the sim- ilarity between the reference regions (marked with yellow stars) and all other spatial locations in the feature maps extracted by DINOv2, as shown in Fig. 4. It is evident that TABLE VIII DETECTION PERFORMANCE WITH DIFFERENT VFM S ACROSS MULTIPLE SCENARIOS . VFM xView→DOTA SRSD→DIOR HRRSD→SSDD 1% 5% 10%...
-
[3]
proposed the first UDA approach built on the one-stage detector RetinaNet [31] and introduced an adaptive feature alignment strategy to address the challenge of accurately aligning sparse foreground objects. Zhu et al. [32] incorporated a mean-teacher framework into remote sensing adaptation, using dual detection heads to suppress biased information and r...
-
[4]
were the first to propose a source-free object detector and to develop a self-entropy based approach for setting an appropriate confidence threshold. Zhang et al. [5] constructed multiple prototypes for instance features within a Faster R- CNN detector and leveraged the consistency between teacher- and student-prototype distributions to correct the genera...
-
[5]
The Effect of Different Vision Foundation Models on Feature Extraction: We further evaluate the performance of our method with different VFMs across three cross-domain scenarios, as reported in Table VIII. DINOv3 [69] outperforms DINOv2 in optical imagery, owing to its ability to capture finer-grained feature representations that are particularly ben- efi...
-
[8]
and feature alignment techniques [1], [9]–[11] are generally infeasible. Instead, self-training paradigms based on the mean teacher framework [12] have become the dominant approach. These paradigms consist of teacher and student models with identical structures, where the teacher model generates pseudo- labels to guide the student model in supervised trai...
work page Pith review arXiv 2025
-
[12]
Unbiased teacher for semi-supervised object detection,
Y .-C. Liu, C.-M. Ma, Z. He, C.-W. Kuo, K. Chen, P. Zhang, B. Wu, Z. Kira, and P. Vajda, “Unbiased teacher for semi-supervised object detection,” in Proceedings of the International Conference on Learning Representations, 2021
work page 2021
-
[13]
Ema: A process model of appraisal dynamics,
S. C. Marsella and J. Gratch, “Ema: A process model of appraisal dynamics,” Cognitive Systems Research , vol. 10, no. 1, p. 70–90, Mar 2009
work page 2009
-
[14]
Periodically exchange teacher-student for source-free object detection,
Q. Liu, L. Lin, Z. Shen, and Z. Yang, “Periodically exchange teacher-student for source-free object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 6414–6424
work page 2023
Show all 63 references
-
[15]
Dynamic retraining-updating mean teacher for source-free object detection,
T. L. B. Khanh, H.-H. Nguyen, L. H. Pham, D. N.-N. Tran, and J. W. Jeon, “Dynamic retraining-updating mean teacher for source-free object detection,” in Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 328–344
2024
-
[16]
Dinov2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[17]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026
2023
-
[18]
Grounding dino: Marrying dino with grounded pre- training for open-set object detection,
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su et al. , “Grounding dino: Marrying dino with grounded pre- training for open-set object detection,” in Proceedings of the European Conference on Computer Vision . Springer, 2024, pp. 38–55
2024
-
[19]
Remoteclip: A vision language foundation model for remote sensing,
F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1– 16, 2024
2024
-
[20]
Some methods for classification and analysis of multi- variate observations,
J. MacQueen, “Some methods for classification and analysis of multi- variate observations,” in Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics , vol. 5. University of California press, 1967, pp. 281–298
1967
-
[21]
Sinkhorn distances: Lightspeed computation of optimal transport,
M. Cuturi, “Sinkhorn distances: Lightspeed computation of optimal transport,” Advances in Neural Information Processing Systems , vol. 26, 2013
2013
-
[22]
xview: Objects in context in overhead imagery,
D. Lam, R. Kuzma, K. McGee, S. Dooley, M. Laielli, M. Klaric, Y . Bulatov, and B. McCord, “xview: Objects in context in overhead imagery,” arXiv preprint arXiv:1802.07856 , 2018
2018 arXiv
-
[23]
Dota: A large-scale dataset for object detection in aerial images,
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 3974–3983
2018
-
[24]
Object detection in optical remote sensing images: A survey and a new benchmark,
K. Li, G. Wan, G. Cheng, L. Meng, and J. Han, “Object detection in optical remote sensing images: A survey and a new benchmark,” ISPRS Journal of Photogrammetry and Remote Sensing , p. 296–307, Jan 2020
2020
-
[25]
Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,
Y . Zhang, Y . Yuan, Y . Feng, and X. Lu, “Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,” IEEE Transactions on Geoscience and Remote Sens- ing, vol. 57, no. 8, pp. 5535–5548, 2019
2019
-
[26]
Ship detection in sar images based on an improved faster r-cnn,
J. Li, C. Qu, and J. Shao, “Ship detection in sar images based on an improved faster r-cnn,” in Proceedings of the 2017 SAR in Big Data Era: Models, Methods and Applications . IEEE, 2017, pp. 1–6
2017
-
[27]
Faster r-cnn: Towards real-time object detection with region proposal networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 6, pp. 1137– 1149, 2017
2017
-
[28]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proceedings of the European Conference on Computer Vision. Springer, 2020, pp. 213–229
2020
-
[29]
Exploring sequence feature alignment for domain adaptive detection transformers,
W. Wang, Y . Cao, J. Zhang, F. He, Z.-J. Zha, Y . Wen, and D. Tao, “Exploring sequence feature alignment for domain adaptive detection transformers,” in Proceedings of the ACM International Conference on Multimedia, 2021, pp. 1730–1738
2021
-
[30]
Da- detr: Domain adaptive detection transformer with information fusion,
J. Zhang, J. Huang, Z. Luo, G. Zhang, X. Zhang, and S. Lu, “Da- detr: Domain adaptive detection transformer with information fusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23 787–23 798
2023
-
[31]
Focal loss for dense object detection,
T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal loss for dense object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2017, pp. 2980–2988
2017
-
[32]
Dualda-net: Dual-head rectification for cross-domain object detection of remote sensing,
Y . Zhu, X. Sun, W. Diao, H. Wei, and K. Fu, “Dualda-net: Dual-head rectification for cross-domain object detection of remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–16, 2023
2023
-
[33]
Remote sensing teacher: Cross-domain detection transformer with learnable frequency- enhanced feature alignment in remote sensing imagery,
J. Han, W. Yang, Y . Wang, L. Chen, and Z. Luo, “Remote sensing teacher: Cross-domain detection transformer with learnable frequency- enhanced feature alignment in remote sensing imagery,” IEEE Trans- actions on Geoscience and Remote Sensing , vol. 62, no. 5619814, pp. 1–14, 2024
2024
-
[34]
Balanced teacher for source-free ob- ject detection,
J. Deng, W. Li, and L. Duan, “Balanced teacher for source-free ob- ject detection,” IEEE Transactions on Circuits and Systems for Video Technology, 2024
2024
-
[35]
Dense learning based semi-supervised object detection,
B. Chen, P. Li, X. Chen, B. Wang, L. Zhang, and X.-S. Hua, “Dense learning based semi-supervised object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4815–4824
2022
-
[36]
Dense teacher: Dense pseudo-labels for semi-supervised object detection,
H. Zhou, Z. Ge, S. Liu, W. Mao, Z. Li, H. Yu, and J. Sun, “Dense teacher: Dense pseudo-labels for semi-supervised object detection,” in Proceedings of the European Conference on Computer Vision. Springer, 2022, pp. 35–50
2022
-
[37]
Efficient non-maximum suppression,
A. Neubeck and L. Van Gool, “Efficient non-maximum suppression,” in Proceedings of the International Conference on Pattern Recognition , 2006, pp. 850–855
2006
-
[38]
End-to-end semi-supervised object detection with soft teacher,
M. Xu, Z. Zhang, H. Hu, J. Wang, L. Wang, F. Wei, X. Bai, and Z. Liu, “End-to-end semi-supervised object detection with soft teacher,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3060–3069
2021
-
[39]
Dual teacher: Improv- ing the reliability of pseudo labels for semi-supervised oriented object detection,
Z. Fang, J. Ren, J. Zheng, R. Chen, and H. Zhao, “Dual teacher: Improv- ing the reliability of pseudo labels for semi-supervised oriented object detection,” IEEE Transactions on Geoscience and Remote Sensing, 2024
2024
-
[40]
Minimizing sample redundancy for label-efficient object detection in aerial images,
R. Zhang, C. Xu, H. Zhu, F. Xu, W. Yang, H. Zhang, and G.-S. Xia, “Minimizing sample redundancy for label-efficient object detection in aerial images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 63, pp. 1–14, 2025
2025
-
[41]
Omni-detr: Omni-supervised object detection with transformers,
P. Wang, Z. Cai, H. Yang, G. Swaminathan, N. Vasconcelos, B. Schiele, and S. Soatto, “Omni-detr: Omni-supervised object detection with transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9367–9376
2022
-
[42]
Semi-detr: Semi-supervised object detection with detection transformers,
J. Zhang, X. Lin, W. Zhang, K. Wang, X. Tan, J. Han, E. Ding, J. Wang, and G. Li, “Semi-detr: Semi-supervised object detection with detection transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 809–23 818
2023
-
[43]
Sparse semi-detr: sparse learnable queries for semi-supervised object detection,
T. Shehzadi, K. A. Hashmi, D. Stricker, and M. Z. Afzal, “Sparse semi-detr: sparse learnable queries for semi-supervised object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5840–5850
2024
-
[44]
Detect everything with few examples,
X. Zhang, Y . Liu, Y . Wang, and A. Boularias, “Detect everything with few examples,” arXiv preprint arXiv:2309.12969 , 2023
2023 arXiv
-
[45]
Cross-domain few-shot object detection via enhanced open-set object detector,
Y . Fu, Y . Wang, Y . Pan, L. Huai, X. Qiu, Z. Shangguan, T. Liu, Y . Fu, L. Van Gool, and X. Jiang, “Cross-domain few-shot object detection via enhanced open-set object detector,” in Proceedings of the European Conference on Computer Vision . Springer, 2024, pp. 247–264
2024
-
[46]
Large self-supervised models bridge the gap in domain adaptive object detection,
M.-A. Lavoie, A. Mahmoud, and S. L. Waslander, “Large self-supervised models bridge the gap in domain adaptive object detection,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 4692–4702
2025
-
[47]
Frozen- detr: Enhancing detr with image understanding from frozen foundation models,
S. Fu, J. Yan, Q. Yang, X. Wei, X. Xie, and W.-S. Zheng, “Frozen- detr: Enhancing detr with image understanding from frozen foundation models,” arXiv preprint arXiv:2410.19635 , 2024
2024 arXiv
-
[48]
Good: Towards domain generalized oriented object detection,
Q. Bi, B. Zhou, J. Yi, W. Ji, H. Zhan, and G.-S. Xia, “Good: Towards domain generalized oriented object detection,” ISPRS Journal of Pho- togrammetry and Remote Sensing , vol. 223, pp. 207–220, 2025
2025
-
[49]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in Proceedings of the International Conference on Machine Learning, pages=8748–8763, ye...
2021
-
[50]
Dino: Detr with improved denoising anchor boxes for end-to- end object detection,
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y . Shum, “Dino: Detr with improved denoising anchor boxes for end-to- end object detection,” in Proceedings of the International Conference on Learning Representations , 2023
2023
-
[51]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proceedings of the International Con...
2021
-
[52]
Exploring robust features for few-shot object detection in satellite 15 imagery,
X. Bou, G. Facciolo, R. G. V on Gioi, J.-M. Morel, and T. Ehret, “Exploring robust features for few-shot object detection in satellite 15 imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 430–439
2024
-
[53]
Segment any change,
Z. Zheng, Y . Zhong, L. Zhang, and S. Ermon, “Segment any change,” Advances in Neural Information Processing Systems , vol. 37, pp. 81 204–81 224, 2024
2024
-
[54]
Evaluating general purpose vision foundation models for medical image analysis: An experimental study of dinov2 on radiology benchmarks,
M. Baharoon, W. Qureshi, J. Ouyang, Y . Xu, A. Aljouie, and W. Peng, “Evaluating general purpose vision foundation models for medical image analysis: An experimental study of dinov2 on radiology benchmarks,” arXiv preprint arXiv:2312.02366 , 2023
2023 arXiv
-
[55]
Dino-reg: General purpose image encoder for training-free multi-modal deformable medical image registration,
X. Song, X. Xu, and P. Yan, “Dino-reg: General purpose image encoder for training-free multi-modal deformable medical image registration,” in International Conference on Medical Image Computing and Computer- Assisted Intervention. Springer, 2024, pp. 608–617
2024
-
[56]
Mask r-cnn,
K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 2961–2969
2017
-
[57]
Datr: Unsupervised domain adaptive detection transformer with dataset-level adaptation and prototypical alignment,
L. Chen, J. Han, and Y . Wang, “Datr: Unsupervised domain adaptive detection transformer with dataset-level adaptation and prototypical alignment,” IEEE Transactions on Image Processing , vol. 34, pp. 982– 994, 2025
2025
-
[58]
Every pixel matters: Center-aware feature alignment for domain adaptive object detector,
C. Hsu, Y .-H. Tsai, Y .-Y . Lin, and M.-H. Yang, “Every pixel matters: Center-aware feature alignment for domain adaptive object detector,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 733–748
2020
-
[59]
Instance relation graph guided source- free domain adaptive object detection,
V . VS, P. Oza, and V . M. Patel, “Instance relation graph guided source- free domain adaptive object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 3520–3530
2023
-
[60]
Enhanc- ing source-free domain adaptive object detection with low-confidence pseudo label distillation,
I. Yoon, H. Kwon, J. Kim, J. Park, H. Jang, and K. Sohn, “Enhanc- ing source-free domain adaptive object detection with low-confidence pseudo label distillation,” in Proceedings of the European Conference on Computer Vision . Springer, 2024, pp. 337–353
2024
-
[61]
Unbiased teacher v2: Semi-supervised object detection for anchor-free and anchor-based detectors,
Y .-C. Liu, C.-Y . Ma, and Z. Kira, “Unbiased teacher v2: Semi-supervised object detection for anchor-free and anchor-based detectors,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9819–9828
2022
-
[62]
Multi-clue consistency learning to bridge gaps between general and oriented object in semi-supervised detection,
C. Wang, C. Xu, X. Li, Y . Li, X. Guo, Z. Gu, and Z. Cui, “Multi-clue consistency learning to bridge gaps between general and oriented object in semi-supervised detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 7, 2025, pp. 7582–7590
2025
-
[63]
Rareplanes: Synthetic data takes flight,
J. Shermeyer, T. Hossler, A. V . Etten, D. Hogan, R. Lewis, and D. Kim, “Rareplanes: Synthetic data takes flight,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision , Jan 2021. [Online]. Available: http://dx.doi.org/10.1109/wacv48630.2021.00025
2021
-
[64]
Unsupervised domain adaptation for remote- sensing vehicle detection using domain-specific channel recalibration,
W. Liu, J. Liu, and B. Luo, “Unsupervised domain adaptation for remote- sensing vehicle detection using domain-specific channel recalibration,” IEEE Geoscience and Remote Sensing Letters , vol. 20, pp. 1–5, 2023
2023
-
[65]
Hierarchical similarity alignment for domain adaptive ship detection in sar images,
J. Zhang, S. Li, Y . Dong, B. Pan, and Z. Shi, “Hierarchical similarity alignment for domain adaptive ship detection in sar images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2022
2022
-
[66]
Fsda-detr: Few- shot domain-adaptive object detection transformer in remote sensing im- agery,
B. Yang, J. Han, X. Hou, D. Zhou, W. Liu, and F. Bi, “Fsda-detr: Few- shot domain-adaptive object detection transformer in remote sensing im- agery,” IEEE Transactions on Geoscience and Remote Sensing , vol. 63, pp. 1–16, 2025
2025
-
[67]
Ima- genet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Ima- genet: A large-scale hierarchical image database,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
-
[68]
Adam: A method for stochastic optimization,
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 , 2014
2014 arXiv
-
[69]
Sim ´eoni, H
O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa et al., “Dinov3,” arXiv preprint arXiv:2508.10104 , 2025
2025 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.