Pith. sign in

REVIEW 3 major objections 1 cited by

Towards integrated sensors for optimized OCT with undetected photons

T0 review · 3 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper claims that integrated Ti:LiNbO3 waveguides can perform OCT with undetected photons at 28 μm axial resolution, and that the induced-coherence scheme is better suited to chip-scale sensors than the SU(1,1) scheme.

desk verdict A credible integrated-quantum-OCT result that I can't verify: the supplied full text is an unrelated CS paper. read the letter →

arxiv 2508.05320 v1 pith:7NYJJB57 submitted 2025-08-07 quant-ph

classification quant-ph
keywords opticalcoherencetomographyundetectedphotonsinducedSU(11)interferometernonlinearwaveguidestitanium-dopedlithiumniobatequantumimagingaxialresolution
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to show that optical coherence tomography with undetected photons—an imaging technique in which the light that touches the sample is never directly detected, its information being read out through a correlated partner photon—can be miniaturized onto integrated nonlinear waveguides. Using titanium-doped lithium niobate (Ti:LiNbO3) waveguides, the authors compare the standard SU(1,1) interferometer scheme with an induced-coherence scheme. They find that the induced-coherence scheme is better suited to integrated systems, and after optimizing the pump gain in both schemes they report axial resolution as low as 28 μm. A sympathetic reader would care because a practical sensor needs both high axial resolution and high brightness, and integrated nonlinear waveguides promise to deliver both in a compact, chip-based device.

What carries the argument

The key objects are two nonlinear interferometer configurations built on Ti:LiNbO3 waveguides. In the SU(1,1) scheme, the nonlinear medium acts as the gain element and both signal and idler photons take part in the interferometer. In the induced-coherence scheme, the sample sits in the path of one photon of a pair, and the readout photon acquires information about the sample only through the coherence induced by path indistinguishability. The waveguides supply broadband photon pairs, whose bandwidth sets the axial resolution and whose brightness sets the achievable signal-to-noise ratio; pump-gain optimization balances these quantities.

What would settle it

Make a second Ti:LiNbO3 device from an independent fabrication run and, at matched pump gain, measure the axial resolution and brightness of the SU(1,1) and induced-coherence schemes on the same target. If the induced-coherence scheme no longer outperforms SU(1,1), or the axial resolution degrades well above 28 μm, the platform-level claim is falsified.

Watch

Extended reading notes

Core claim

The central discovery, on the paper's own terms, is that integrated Ti:LiNbO3 waveguides can support OCT with undetected photons in both standard nonlinear interferometer configurations, and that the induced-coherence scheme is the more appropriate one for integrated sensors. The authors perform pump-gain optimization in each scheme and obtain axial resolution down to 28 μm. This positions the induced-coherence configuration, rather than SU(1,1), as the practical route to miniaturized OCT sensors that measure at a wavelength where detection is difficult while reading out at a convenient wavelength.

Load-bearing premise

The load-bearing premise is that the measured 28 μm axial resolution and the induced-coherence scheme's advantage are properties of the integrated-waveguide platform in general, not just of the particular crystal sample and pump settings used.

Editorial extensions

If this is right

  • Undetected-photon OCT can be implemented on a chip, replacing bulk nonlinear crystals with Ti:LiNbO3 waveguides.
  • Future integrated OCT sensors should start from the induced-coherence scheme rather than SU(1,1).
  • A resolution around 28 μm is compatible with a bright, pump-optimized source, making the integrated approach relevant for practical imaging.
  • Because readout uses the detected partner photon, the sample can be probed at wavelengths where efficient detectors do not exist.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 28 μm result reflects the platform rather than one sample, dispersion engineering of the waveguides could push the resolution further, since axial resolution in OCT scales inversely with available spectral bandwidth.
  • The induced-coherence scheme's advantage may become more pronounced as devices shrink, because it avoids some of the loss and complexity of running the SU(1,1) configuration on chip.
  • A natural next experiment is to use the same integrated source for mid-infrared OCT or spectroscopy, where the undetected-photon trick matters most.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The abstract describes an experimental quantum-optics study: nonlinear Ti:LiNbO3 waveguides are used for optical coherence tomography (OCT) with undetected photons, comparing the SU(1,1) and induced-coherence schemes, with pump-gain optimization and a claimed axial resolution as low as 28 μm. The stated contribution is that the induced-coherence scheme is better suited to integrated systems. However, the body of the manuscript supplied for review is an unrelated computer-science paper, mKG-RAG (arXiv:2508.05318v2), on multimodal knowledge graphs for visual question answering. No quantum-optics derivation, experimental setup, data, or analysis appears. The abstract's claims cannot be checked against any supporting content.

Significance. If the abstract's findings are correct, they would be significant: demonstrating OCT with undetected photons at 28 μm axial resolution in integrated Ti:LiNbO3 waveguides, and showing that the induced-coherence scheme outperforms SU(1,1) for integrated sensors, would be a useful step toward miniaturized quantum-enabled OCT systems. The comparative benchmarking between the two schemes is practically valuable. However, in the submitted manuscript these claims are entirely unsupported. There are no machine-checked proofs, reproducible code, data tables, or experimental artifacts to credit, so the significance cannot be assessed beyond the abstract's assertion.

major comments (3)
  1. [Full Text (entire body)] The supplied full text is an unrelated paper, mKG-RAG (arXiv:2508.05318v2), on retrieval-augmented generation for visual question answering. None of the OCT-related content from the abstract appears in the body: there is no description of the Ti:LiNbO3 waveguides, no pump-gain optimization procedure, no experimental layout for either the SU(1,1) or induced-coherence scheme, and no measurement of axial resolution. This is a complete absence of support for the central claims, not a local gap.
  2. [Abstract, last sentence] The abstract claims an axial resolution as low as 28 μm and states that the induced-coherence scheme is 'better suited' for integrated systems. Because the supporting text is absent, I cannot determine whether 28 μm is a representative platform result or a best-case single-sample value, nor whether the comparison between schemes is controlled (e.g., matched pump gain, detection efficiency, waveguide parameters).
  3. [Full Text (no quantum-optics sections)] There are no equations, data tables, or figures relevant to OCT with undetected photons. As a result, the pump-gain optimization and the SU(1,1)-vs-induced-coherence comparison cannot be verified at any level of technical detail. This precludes a soundness assessment.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable: the supplied full text is an unrelated mKG-RAG paper, so the OCT derivation chain cannot be audited, but the abstract itself shows no circular structure.

full rationale

The submitted abstract describes an experimental comparison of two established quantum-optics schemes (SU(1,1) and induced coherence) for OCT with undetected photons in Ti:LiNbO3 waveguides, reporting pump-gain optimization and an axial resolution as low as 28 μm. The supplied full text, however, is an entirely different manuscript (mKG-RAG, arXiv:2508.05318v2), so there is no derivation chain, equations, or experimental details from the OCT paper to examine. Within the abstract alone, the claims are empirical benchmark results against external measurements, not derived from fitted parameters, self-citations, or definitions that presuppose the conclusion. The 28 μm resolution and the comparative judgment about integrated suitability are unverifiable from the provided material, but that is a completeness/integrity limitation, not circularity. No step in the visible text reduces to its own input by construction, and no load-bearing self-citation is in evidence. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The central experimental conclusion rests on established nonlinear interferometer physics. The only adjustable experimental parameter mentioned in the abstract is pump gain, which is optimized in the measurement rather than introduced as a theoretical assumption. No new entities are postulated.

free parameters (1)
  • pump gain = not stated (optimized in experiment)
    Optimized in the experiment to maximize brightness; the abstract does not state its value. It is an experimental control, not a fitted model constant.
assumptions (2)
  • domain assumption SU(1,1) interferometer physics
    The comparison assumes standard nonlinear interferometer behavior as known from prior literature; the abstract does not derive it.
  • domain assumption Induced coherence without induced emission
    Assumes the induced coherence scheme behaves as in earlier demonstrations; the abstract relies on this established quantum-optics effect.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards integrated sensors for optimized OCT with undetected photons." pith.science (2026). https://pith.science/paper/7NYJJB57

@misc{pith2026250805320,
  author       = {Pith},
  title        = {Pith review of: Towards integrated sensors for optimized OCT with undetected photons},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7NYJJB57}},
  note         = {Machine review of arXiv:2508.05320}
}
abstract

The development of practical sensors for optical coherence tomography (OCT) with undetected photons requires miniaturization via integration. To be practical, these sensors must exhibit a large spectral bandwidth and a high brightness, which are linked to a high axial resolution and a sufficient signal to noise ratio, respectively. Here, we combine these requirements in a scheme for OCT-measurements with undetected photons based on nonlinear Ti:LiNbO$_3$ waveguides. We investigate the performance benchmarks of the commonly used SU(1,1) scheme in comparison to an induced coherence scheme and find that the latter is actually better suited when implementing measurements with undetected photons in integrated systems. In both schemes, we perform pump gain optimization and OCT measurements with undetected photons with an axial resolution as low as 28 $\mathrm{\mu}$m.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Below-shot-noise capacity in phase estimation using nonlinear interferometers

    quant-ph 2026-01 conditional novelty 6.0 of 10

    For phase estimation with intensity measurements in lossy nonlinear interferometers, differential detection in a Mandel-type setup is robust and shot-noise-limited, while a Yurke-type setup only achieves sub-shot-nois...

Reference graph

Works this paper leans on

75 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, and Amit Bahree. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.arXiv preprint arXiv:2404.14219 (2024)

  2. [2]

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision. 2425–2433

  3. [3]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  4. [4]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901

  5. [5]

    Chenyang Bu, Guojie Chang, Zihao Chen, Cunyuan Dang, Zhize Wu, Yi He, and Xindong Wu. 2025. Query-Driven Multimodal GraphRAG: Dynamic Local Knowl- edge Graph Construction for Online Reasoning. InFindings of the Association for Computational Linguistics: ACL 2025. 21360–21380

  6. [6]

    Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 291–305

  7. [7]

    Davide Caffagni, Federico Cocchi, Nicholas Moratelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024. Wiki-llava: Hierarchical retrieval- augmented generation for multimodal llms. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition. 1818–1826

  8. [8]

    Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023. Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 14948–14968

Show all 75 references
  1. [9]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al . 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on compu...

  2. [10]

    Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2025. LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.arXiv preprint arXiv:2503.15621(2025)

  3. [11]

    Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2025. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  4. [12]

    Xinnan Dai, Haohao Qu, Yifei Shen, Bohang Zhang, Qihao Wen, Wenqi Fan, Dong- sheng Li, Jiliang Tang, and Caihua Shan. 2025. How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension. InThe Thirteenth International Conference on Learnin...

  5. [13]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130(2024)

  6. [14]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ...

  7. [15]

    Wenqi Fan, Yao Ma, Qing Li, Yuan He, Eric Zhao, Jiliang Tang, and Dawei Yin

  8. [16]

    Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018. Iqa: Visual question answering in interactive environments. InProceedings of the IEEE conference on computer vision and pattern recognition. 4089–4098

  9. [17]

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh

  10. [18]

    Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, and Chao Huang. 2025. RAG- Anything: All-in-One RAG Framework.arXiv preprint arXiv:2510.12323(2025)

  11. [19]

    Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2025. LightRAG: Simple and Fast Retrieval-Augmented Generation. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

  12. [20]

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual representation learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9729–9738

  13. [21]

    Wenjue He, Zheng Zhang, Yongyong Chen, and Jie Wen. 2023. Structured anchor- inferred graph learning for universal incomplete multi-view clustering.World Wide Web26, 1 (2023), 375–399

  14. [22]

    Wen-Jue He, Xiaofeng Zhu, and Zheng Zhang. 2026. Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 17463–17471

  15. [23]

    Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Navigli, Sebastian Neumaier, et al. 2021. Knowledge graphs.ACM Computing Surveys (Csur)54, 4 (2021), 1–37

  16. [24]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.ICLR1, 2 (2022), 3

  17. [25]

    Jiani Huang, Shijie Wang, Liangbo Ning, Wenqi Fan, and Qing Li. 2026. ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning.arXiv preprint arXiv:2604.07851(2026)

  18. [26]

    Zhijing Huang, Wen-Jue He, Baotian Hu, and Zheng Zhang. 2026. Grading- Inspired Complementary Enhancing for Multimodal Sentiment Analysis.Infor- mation Fusion(2026), 104174

  19. [27]

    Jinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee, and Seunghyun Park

  20. [28]

    Zhuohang Jiang, Pangjing Wu, Ziran Liang, Peter Q Chen, Xu Yuan, Ye Jia, Jiancheng Tu, Chen Li, Peter HF Ng, and Qing Li. 2025. Hibench: Benchmarking llms capability on hierarchical structure reasoning. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and...

  21. [29]

    Zhuohang Jiang, Pangjing Wu, Xu Yuan, Wenqi Fan, and Qing Li. 2025. QA- Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering.arXiv preprint arXiv:2508.05197(2025)

  22. [30]

    Zhuohang Jiang, Xu Yuan, Haohao Qu, Shanru Lin, Kanglong Liu, Wenqi Fan, and Qing Li. 2026. SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses.arXiv preprint arXiv:2602.22683(2026)

  23. [31]

    Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs.IEEE Transactions on Big Data7, 3 (2019), 535–547

  24. [32]

    Junlin Lee, Yequan Wang, Jing Li, and Min Zhang. 2024. Multimodal Reasoning with Multimodal Knowledge Graph. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10767– 10782

  25. [33]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational conference on machine learning. PMLR, 19730–19742

  26. [34]

    Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022. Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning.Advances in Neural Information Processing Systems35 (2022), 17612–17625

  27. [35]

    Weizhe Lin and Bill Byrne. 2022. Retrieval Augmented Visual Question Answering with Outside Knowledge. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 11238–11254

  28. [36]

    Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, and Bill Byrne. 2023. Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering. InAdvances in Neural Information Processing Systems, Vol. 36. 22820–22840

  29. [37]

    Zhihong Lin, Donghao Zhang, Qingyi Tao, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. 2023. Medical visual question answering: A survey.Artificial Intelligence in Medicine143 (2023), 102611

  30. [38]

    Chengliang Liu, Liangbo Ning, Yujuan Ding, and Wenqi Fan. 2026. Inference Cost Attacks for Retrieval-Augmented Large Language Models. InProceedings of the ACM Web Conference 2026. 7564–7575

  31. [39]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26296–26306

  32. [40]

    Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019. MMKG: multi-modal knowledge graphs. InThe semantic web: 16th international conference, ESWC 2019, portorož, Slovenia, June 2–6, 2019, proceedings 16. Springer, 459–474

  33. [41]

    Haoran Luo, Haihong E, Guanting Chen, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng, Zemin Kuang, Meina Song, Yifan Zhu, and Anh Tuan Luu. 2025. HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation. InThe Thirty-ninth Annual...

  34. [42]

    Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, and Hanjiang Lai. 2025. HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation.arXiv preprint arXiv:2511.15435(2025)

  35. [43]

    LINHAO LUO, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2024. Reason- ing on Graphs: Faithful and Interpretable Large Language Model Reasoning. In The Twelfth International Conference on Learning Representations. ����� ���� ���� ������ ����� ���������� ���� ���������� ��...

  36. [44]

    Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, Cehao Yang, Jiaxin Mao, and Jian Guo. 2025. Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation. In The Thirteenth International Conference on Lear...

  37. [45]

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3195–3204

  38. [46]

    Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023. Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories. In Proceedings of the IEEE/CVF International Co...

  39. [47]

    Meta-AI. 2024. Llama 3.2: Revolutionizing edge AI and vision with open, cus- tomizable models. https://ai.meta.com/blog/llama-3-2-connect-2024-vision- edge-mobile-devices

  40. [48]

    Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar. 2020. Plotqa: Reasoning over scientific plots. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1527–1536

  41. [49]

    Liangbo Ning, Wenqi Fan, and Qing Li. 2025. Retrieval-augmented purifier for robust LLM-empowered recommendation.ACM Transactions on Information Systems(2025)

  42. [50]

    Zach Nussbaum, John Xavier Morris, Andriy Mulyar, and Brandon Duderstadt

  43. [51]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems35 (...

  44. [52]

    Jingyuan Qi, Zhiyang Xu, Rulin Shao, Yang Chen, Jin Di, Yu Cheng, Qifan Wang, and Lifu Huang. 2024. Rora-vlm: Robust retrieval-augmented vision language models.arXiv preprint arXiv:2410.08876(2024)

  45. [53]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...

  46. [54]

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. InAdvances in neural information processing systems, Vol. 28

  47. [55]

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. InEuropean conference on computer vision. Springer, 146–162

  48. [56]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)

  49. [57]

    Xueyao Wan and Hang Yu. 2025. MMGraphRAG: Bridging Vision and Lan- guage with Interpretable Multimodal Knowledge Graphs.arXiv preprint arXiv:2507.20804(2025)

  50. [58]

    P Wang, S Bai, S Tan, S Wang, Z Fan, J Bai, K Chen, X Liu, J Wang, W Ge, et al

  51. [59]

    Shijie Wang, Wenqi Fan, Yue Feng, Lin Shanru, Xinyu Ma, Shuaiqiang Wang, and Dawei Yin. 2025. Knowledge graph retrieval-augmented generation for llm-based recommendation. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long ...

  52. [60]

    Shijie Wang, Jiani Huang, Zhikai Chen, Yu Song, Wenzhuo Tang, Haitao Mao, Wenqi Fan, Hui Liu, Xiaorui Liu, Dawei Yin, et al. 2025. Graph machine learning in the era of large language models (llms).ACM Transactions on Intelligent Systems and Technology16, 5 (2025), 1–40

  53. [61]

    Xiaoyang Wang, Yao Ma, Yiqi Wang, Wei Jin, Xin Wang, Jiliang Tang, Caiyan Jia, and Jian Yu. 2020. Traffic flow prediction via spatial temporal graph neural network. InProceedings of the web conference 2020. 1082–1092

  54. [62]

    Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. 2024. Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal under- standing.arXiv preprint arXiv:2412.10302(2024)

  55. [63]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.URL https://arxiv.org/abs/2409.12191(2024)

  56. [64]

    Xu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou, and Jingsong Lan. 2025. Instruction- guided multi-granularity segmentation and captioning with large multimodal model. InProceedings of the AAAI Conference on Artificial Intelligence

  57. [65]

    Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu. 2024. Osprey: Pixel understanding with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 28202–28211

  58. [66]

    Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Yuxuan Zhao, Zehua Xie, et al. 2024. mR2AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA.arXiv preprint arXiv:2411.15041(2024)

  59. [67]

    Zheng Zhang and Wen-Jue He. 2023. Tensorized topological graph learning for generalized incomplete multi-view clustering.Information Fusion100 (2023), 101914

  60. [68]

    Yibin Yan and Weidi Xie. 2024. EchoSight: Advancing Visual-Language Models with Wiki Knowledge. InFindings of the Association for Computational Linguistics: EMNLP 2024. 1538–1551

  61. [69]

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479(2025)

  62. [70]

    Xiangrong Zhu, Yuexiang Xie, Yi Liu, Yaliang Li, and Wei Hu. 2025. Knowledge graph-guided retrieval augmented generation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies...

  63. [73]

    Zheng Zhang, Xu Yuan, Lei Zhu, Jingkuan Song, and Liqiang Nie. 2024. Badcm: Invisible backdoor attack against cross-modal learning.IEEE Transactions on Image Processing33 (2024), 2558–2571

  64. [2017]

    InProceedings of the IEEE conference on computer vision and pattern recognition

    Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 6904–6913

  65. [2019]

    InThe world wide web conference

    Graph neural networks for social recommendation. InThe world wide web conference. 417–426

  66. [2024]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Egtr: Extracting graph from transformer for scene graph generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 24229–24238

  67. [2025]

    Transactions on Machine Learning Research(2025)

    Nomic Embed: Training a Reproducible Long Context Text Embedder. Transactions on Machine Learning Research(2025)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.