REVIEW 3 major objections 1 cited by
Towards integrated sensors for optimized OCT with undetected photons
T0 review · 3 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that integrated Ti:LiNbO3 waveguides can perform OCT with undetected photons at 28 μm axial resolution, and that the induced-coherence scheme is better suited to chip-scale sensors than the SU(1,1) scheme.
desk verdict A credible integrated-quantum-OCT result that I can't verify: the supplied full text is an unrelated CS paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key objects are two nonlinear interferometer configurations built on Ti:LiNbO3 waveguides. In the SU(1,1) scheme, the nonlinear medium acts as the gain element and both signal and idler photons take part in the interferometer. In the induced-coherence scheme, the sample sits in the path of one photon of a pair, and the readout photon acquires information about the sample only through the coherence induced by path indistinguishability. The waveguides supply broadband photon pairs, whose bandwidth sets the axial resolution and whose brightness sets the achievable signal-to-noise ratio; pump-gain optimization balances these quantities.
What would settle it
Make a second Ti:LiNbO3 device from an independent fabrication run and, at matched pump gain, measure the axial resolution and brightness of the SU(1,1) and induced-coherence schemes on the same target. If the induced-coherence scheme no longer outperforms SU(1,1), or the axial resolution degrades well above 28 μm, the platform-level claim is falsified.
Extended reading notes
Core claim
The central discovery, on the paper's own terms, is that integrated Ti:LiNbO3 waveguides can support OCT with undetected photons in both standard nonlinear interferometer configurations, and that the induced-coherence scheme is the more appropriate one for integrated sensors. The authors perform pump-gain optimization in each scheme and obtain axial resolution down to 28 μm. This positions the induced-coherence configuration, rather than SU(1,1), as the practical route to miniaturized OCT sensors that measure at a wavelength where detection is difficult while reading out at a convenient wavelength.
Load-bearing premise
The load-bearing premise is that the measured 28 μm axial resolution and the induced-coherence scheme's advantage are properties of the integrated-waveguide platform in general, not just of the particular crystal sample and pump settings used.
Editorial extensions
If this is right
- Undetected-photon OCT can be implemented on a chip, replacing bulk nonlinear crystals with Ti:LiNbO3 waveguides.
- Future integrated OCT sensors should start from the induced-coherence scheme rather than SU(1,1).
- A resolution around 28 μm is compatible with a bright, pump-optimized source, making the integrated approach relevant for practical imaging.
- Because readout uses the detected partner photon, the sample can be probed at wavelengths where efficient detectors do not exist.
Reading between the lines
- If the 28 μm result reflects the platform rather than one sample, dispersion engineering of the waveguides could push the resolution further, since axial resolution in OCT scales inversely with available spectral bandwidth.
- The induced-coherence scheme's advantage may become more pronounced as devices shrink, because it avoids some of the loss and complexity of running the SU(1,1) configuration on chip.
- A natural next experiment is to use the same integrated source for mid-infrared OCT or spectroscopy, where the undetected-photon trick matters most.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract describes an experimental quantum-optics study: nonlinear Ti:LiNbO3 waveguides are used for optical coherence tomography (OCT) with undetected photons, comparing the SU(1,1) and induced-coherence schemes, with pump-gain optimization and a claimed axial resolution as low as 28 μm. The stated contribution is that the induced-coherence scheme is better suited to integrated systems. However, the body of the manuscript supplied for review is an unrelated computer-science paper, mKG-RAG (arXiv:2508.05318v2), on multimodal knowledge graphs for visual question answering. No quantum-optics derivation, experimental setup, data, or analysis appears. The abstract's claims cannot be checked against any supporting content.
Significance. If the abstract's findings are correct, they would be significant: demonstrating OCT with undetected photons at 28 μm axial resolution in integrated Ti:LiNbO3 waveguides, and showing that the induced-coherence scheme outperforms SU(1,1) for integrated sensors, would be a useful step toward miniaturized quantum-enabled OCT systems. The comparative benchmarking between the two schemes is practically valuable. However, in the submitted manuscript these claims are entirely unsupported. There are no machine-checked proofs, reproducible code, data tables, or experimental artifacts to credit, so the significance cannot be assessed beyond the abstract's assertion.
major comments (3)
- [Full Text (entire body)] The supplied full text is an unrelated paper, mKG-RAG (arXiv:2508.05318v2), on retrieval-augmented generation for visual question answering. None of the OCT-related content from the abstract appears in the body: there is no description of the Ti:LiNbO3 waveguides, no pump-gain optimization procedure, no experimental layout for either the SU(1,1) or induced-coherence scheme, and no measurement of axial resolution. This is a complete absence of support for the central claims, not a local gap.
- [Abstract, last sentence] The abstract claims an axial resolution as low as 28 μm and states that the induced-coherence scheme is 'better suited' for integrated systems. Because the supporting text is absent, I cannot determine whether 28 μm is a representative platform result or a best-case single-sample value, nor whether the comparison between schemes is controlled (e.g., matched pump gain, detection efficiency, waveguide parameters).
- [Full Text (no quantum-optics sections)] There are no equations, data tables, or figures relevant to OCT with undetected photons. As a result, the pump-gain optimization and the SU(1,1)-vs-induced-coherence comparison cannot be verified at any level of technical detail. This precludes a soundness assessment.
Circularity Check
No circularity identifiable: the supplied full text is an unrelated mKG-RAG paper, so the OCT derivation chain cannot be audited, but the abstract itself shows no circular structure.
full rationale
The submitted abstract describes an experimental comparison of two established quantum-optics schemes (SU(1,1) and induced coherence) for OCT with undetected photons in Ti:LiNbO3 waveguides, reporting pump-gain optimization and an axial resolution as low as 28 μm. The supplied full text, however, is an entirely different manuscript (mKG-RAG, arXiv:2508.05318v2), so there is no derivation chain, equations, or experimental details from the OCT paper to examine. Within the abstract alone, the claims are empirical benchmark results against external measurements, not derived from fitted parameters, self-citations, or definitions that presuppose the conclusion. The 28 μm resolution and the comparative judgment about integrated suitability are unverifiable from the provided material, but that is a completeness/integrity limitation, not circularity. No step in the visible text reduces to its own input by construction, and no load-bearing self-citation is in evidence. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- pump gain =
not stated (optimized in experiment)
assumptions (2)
- domain assumption SU(1,1) interferometer physics
- domain assumption Induced coherence without induced emission
Cite this review
Pith. "Pith review of Towards integrated sensors for optimized OCT with undetected photons." pith.science (2026). https://pith.science/paper/7NYJJB57
@misc{pith2026250805320,
author = {Pith},
title = {Pith review of: Towards integrated sensors for optimized OCT with undetected photons},
year = {2026},
howpublished = {\url{https://pith.science/paper/7NYJJB57}},
note = {Machine review of arXiv:2508.05320}
}
abstract
The development of practical sensors for optical coherence tomography (OCT) with undetected photons requires miniaturization via integration. To be practical, these sensors must exhibit a large spectral bandwidth and a high brightness, which are linked to a high axial resolution and a sufficient signal to noise ratio, respectively. Here, we combine these requirements in a scheme for OCT-measurements with undetected photons based on nonlinear Ti:LiNbO$_3$ waveguides. We investigate the performance benchmarks of the commonly used SU(1,1) scheme in comparison to an induced coherence scheme and find that the latter is actually better suited when implementing measurements with undetected photons in integrated systems. In both schemes, we perform pump gain optimization and OCT measurements with undetected photons with an axial resolution as low as 28 $\mathrm{\mu}$m.
Forward citations
Cited by 1 Pith paper
-
Below-shot-noise capacity in phase estimation using nonlinear interferometers
For phase estimation with intensity measurements in lossy nonlinear interferometers, differential detection in a Mandel-type setup is robust and shot-noise-limited, while a Yurke-type setup only achieves sub-shot-nois...
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, and Amit Bahree. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.arXiv preprint arXiv:2404.14219 (2024)
arXiv 2024
-
[2]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision. 2425–2433
2015
-
[3]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
arXiv 2025
-
[4]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901
2020
-
[5]
Chenyang Bu, Guojie Chang, Zihao Chen, Cunyuan Dang, Zhize Wu, Yi He, and Xindong Wu. 2025. Query-Driven Multimodal GraphRAG: Dynamic Local Knowl- edge Graph Construction for Online Reasoning. InFindings of the Association for Computational Linguistics: ACL 2025. 21360–21380
work page 2025
-
[6]
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 291–305
work page 2022
-
[7]
Davide Caffagni, Federico Cocchi, Nicholas Moratelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024. Wiki-llava: Hierarchical retrieval- augmented generation for multimodal llms. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition. 1818–1826
2024
-
[8]
Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023. Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 14948–14968
work page 2023
Show all 75 references
-
[9]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al . 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on compu...
2024
-
[10]
Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2025. LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.arXiv preprint arXiv:2503.15621(2025)
2025 arXiv
-
[11]
Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2025. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2025
-
[12]
Xinnan Dai, Haohao Qu, Yifei Shen, Bohang Zhang, Qihao Wen, Wenqi Fan, Dong- sheng Li, Jiliang Tang, and Caihua Shan. 2025. How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension. InThe Thirteenth International Conference on Learnin...
2025
-
[13]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130(2024)
2024 arXiv
-
[14]
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ...
2024
-
[15]
Wenqi Fan, Yao Ma, Qing Li, Yuan He, Eric Zhao, Jiliang Tang, and Dawei Yin
-
[16]
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018. Iqa: Visual question answering in interactive environments. InProceedings of the IEEE conference on computer vision and pattern recognition. 4089–4098
2018
-
[17]
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh
-
[18]
Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, and Chao Huang. 2025. RAG- Anything: All-in-One RAG Framework.arXiv preprint arXiv:2510.12323(2025)
2025
-
[19]
Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2025. LightRAG: Simple and Fast Retrieval-Augmented Generation. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
2025
-
[20]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual representation learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9729–9738
2020
-
[21]
Wenjue He, Zheng Zhang, Yongyong Chen, and Jie Wen. 2023. Structured anchor- inferred graph learning for universal incomplete multi-view clustering.World Wide Web26, 1 (2023), 375–399
2023
-
[22]
Wen-Jue He, Xiaofeng Zhu, and Zheng Zhang. 2026. Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 17463–17471
2026
-
[23]
Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Navigli, Sebastian Neumaier, et al. 2021. Knowledge graphs.ACM Computing Surveys (Csur)54, 4 (2021), 1–37
2021
-
[24]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.ICLR1, 2 (2022), 3
2022
-
[25]
Jiani Huang, Shijie Wang, Liangbo Ning, Wenqi Fan, and Qing Li. 2026. ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning.arXiv preprint arXiv:2604.07851(2026)
2026 arXiv
-
[26]
Zhijing Huang, Wen-Jue He, Baotian Hu, and Zheng Zhang. 2026. Grading- Inspired Complementary Enhancing for Multimodal Sentiment Analysis.Infor- mation Fusion(2026), 104174
2026
-
[27]
Jinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee, and Seunghyun Park
-
[28]
Zhuohang Jiang, Pangjing Wu, Ziran Liang, Peter Q Chen, Xu Yuan, Ye Jia, Jiancheng Tu, Chen Li, Peter HF Ng, and Qing Li. 2025. Hibench: Benchmarking llms capability on hierarchical structure reasoning. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and...
2025
-
[29]
Zhuohang Jiang, Pangjing Wu, Xu Yuan, Wenqi Fan, and Qing Li. 2025. QA- Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering.arXiv preprint arXiv:2508.05197(2025)
2025
-
[30]
Zhuohang Jiang, Xu Yuan, Haohao Qu, Shanru Lin, Kanglong Liu, Wenqi Fan, and Qing Li. 2026. SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses.arXiv preprint arXiv:2602.22683(2026)
2026 arXiv
-
[31]
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs.IEEE Transactions on Big Data7, 3 (2019), 535–547
2019
-
[32]
Junlin Lee, Yequan Wang, Jing Li, and Min Zhang. 2024. Multimodal Reasoning with Multimodal Knowledge Graph. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10767– 10782
2024
-
[33]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational conference on machine learning. PMLR, 19730–19742
2023
-
[34]
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022. Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning.Advances in Neural Information Processing Systems35 (2022), 17612–17625
2022
-
[35]
Weizhe Lin and Bill Byrne. 2022. Retrieval Augmented Visual Question Answering with Outside Knowledge. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 11238–11254
2022
-
[36]
Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, and Bill Byrne. 2023. Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering. InAdvances in Neural Information Processing Systems, Vol. 36. 22820–22840
2023
-
[37]
Zhihong Lin, Donghao Zhang, Qingyi Tao, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. 2023. Medical visual question answering: A survey.Artificial Intelligence in Medicine143 (2023), 102611
2023
-
[38]
Chengliang Liu, Liangbo Ning, Yujuan Ding, and Wenqi Fan. 2026. Inference Cost Attacks for Retrieval-Augmented Large Language Models. InProceedings of the ACM Web Conference 2026. 7564–7575
2026
-
[39]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26296–26306
2024
-
[40]
Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019. MMKG: multi-modal knowledge graphs. InThe semantic web: 16th international conference, ESWC 2019, portorož, Slovenia, June 2–6, 2019, proceedings 16. Springer, 459–474
2019
-
[41]
Haoran Luo, Haihong E, Guanting Chen, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng, Zemin Kuang, Meina Song, Yifan Zhu, and Anh Tuan Luu. 2025. HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation. InThe Thirty-ninth Annual...
2025
-
[42]
Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, and Hanjiang Lai. 2025. HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation.arXiv preprint arXiv:2511.15435(2025)
2025
-
[43]
LINHAO LUO, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2024. Reason- ing on Graphs: Faithful and Interpretable Large Language Model Reasoning. In The Twelfth International Conference on Learning Representations. ����� ���� ���� ������ ����� ���������� ���� ���������� ��...
2024
-
[44]
Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, Cehao Yang, Jiaxin Mao, and Jian Guo. 2025. Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation. In The Thirteenth International Conference on Lear...
2025
-
[45]
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3195–3204
2019
-
[46]
Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023. Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories. In Proceedings of the IEEE/CVF International Co...
2023
-
[47]
Meta-AI. 2024. Llama 3.2: Revolutionizing edge AI and vision with open, cus- tomizable models. https://ai.meta.com/blog/llama-3-2-connect-2024-vision- edge-mobile-devices
2024
-
[48]
Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar. 2020. Plotqa: Reasoning over scientific plots. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1527–1536
2020
-
[49]
Liangbo Ning, Wenqi Fan, and Qing Li. 2025. Retrieval-augmented purifier for robust LLM-empowered recommendation.ACM Transactions on Information Systems(2025)
2025
-
[50]
Zach Nussbaum, John Xavier Morris, Andriy Mulyar, and Brandon Duderstadt
-
[51]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems35 (...
2022
-
[52]
Jingyuan Qi, Zhiyang Xu, Rulin Shao, Yang Chen, Jin Di, Yu Cheng, Qifan Wang, and Lifu Huang. 2024. Rora-vlm: Robust retrieval-augmented vision language models.arXiv preprint arXiv:2410.08876(2024)
2024 arXiv
-
[53]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[54]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. InAdvances in neural information processing systems, Vol. 28
2015
-
[55]
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. InEuropean conference on computer vision. Springer, 146–162
2022
-
[56]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)
2023 arXiv
-
[57]
Xueyao Wan and Hang Yu. 2025. MMGraphRAG: Bridging Vision and Lan- guage with Interpretable Multimodal Knowledge Graphs.arXiv preprint arXiv:2507.20804(2025)
2025 arXiv
-
[58]
P Wang, S Bai, S Tan, S Wang, Z Fan, J Bai, K Chen, X Liu, J Wang, W Ge, et al
-
[59]
Shijie Wang, Wenqi Fan, Yue Feng, Lin Shanru, Xinyu Ma, Shuaiqiang Wang, and Dawei Yin. 2025. Knowledge graph retrieval-augmented generation for llm-based recommendation. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long ...
2025
-
[60]
Shijie Wang, Jiani Huang, Zhikai Chen, Yu Song, Wenzhuo Tang, Haitao Mao, Wenqi Fan, Hui Liu, Xiaorui Liu, Dawei Yin, et al. 2025. Graph machine learning in the era of large language models (llms).ACM Transactions on Intelligent Systems and Technology16, 5 (2025), 1–40
2025
-
[61]
Xiaoyang Wang, Yao Ma, Yiqi Wang, Wei Jin, Xin Wang, Jiliang Tang, Caiyan Jia, and Jian Yu. 2020. Traffic flow prediction via spatial temporal graph neural network. InProceedings of the web conference 2020. 1082–1092
2020
-
[62]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. 2024. Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal under- standing.arXiv preprint arXiv:2412.10302(2024)
2024 arXiv
-
[63]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.URL https://arxiv.org/abs/2409.12191(2024)
2024 arXiv
-
[64]
Xu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou, and Jingsong Lan. 2025. Instruction- guided multi-granularity segmentation and captioning with large multimodal model. InProceedings of the AAAI Conference on Artificial Intelligence
2025
-
[65]
Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu. 2024. Osprey: Pixel understanding with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 28202–28211
2024
-
[66]
Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Yuxuan Zhao, Zehua Xie, et al. 2024. mR2AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA.arXiv preprint arXiv:2411.15041(2024)
2024 arXiv
-
[67]
Zheng Zhang and Wen-Jue He. 2023. Tensorized topological graph learning for generalized incomplete multi-view clustering.Information Fusion100 (2023), 101914
2023
-
[68]
Yibin Yan and Weidi Xie. 2024. EchoSight: Advancing Visual-Language Models with Wiki Knowledge. InFindings of the Association for Computational Linguistics: EMNLP 2024. 1538–1551
2024
-
[69]
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479(2025)
2025 arXiv
-
[70]
Xiangrong Zhu, Yuexiang Xie, Yi Liu, Yaliang Li, and Wei Hu. 2025. Knowledge graph-guided retrieval augmented generation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies...
2025
-
[73]
Zheng Zhang, Xu Yuan, Lei Zhu, Jingkuan Song, and Liqiang Nie. 2024. Badcm: Invisible backdoor attack against cross-modal learning.IEEE Transactions on Image Processing33 (2024), 2558–2571
2024
-
[2017]
InProceedings of the IEEE conference on computer vision and pattern recognition
Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 6904–6913
-
[2019]
InThe world wide web conference
Graph neural networks for social recommendation. InThe world wide web conference. 417–426
-
[2024]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Egtr: Extracting graph from transformer for scene graph generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 24229–24238
-
[2025]
Transactions on Machine Learning Research(2025)
Nomic Embed: Training a Reproducible Long Context Text Embedder. Transactions on Machine Learning Research(2025)
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.