REVIEW 4 major objections 5 minor 89 references
DBMS-LLM Integration Strategies in Industrial and Business Applications: Current Status and Future Challenges
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper argues that DBMS-LLM integration in industry falls into five architectural patterns, each with distinct coupling and trade-offs.
desk verdict Useful survey of DBMS-LLM integration, but the five-way taxonomy is a list of overlapping perspectives, not a principled partition; still deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a five-category taxonomy built on three dimensions: the purpose of integration, the system layer where integration occurs, and the degree of coupling. The taxonomy is supported by a comparison table of strengths and weaknesses and a feature-scoring table that rates each pattern as strong, moderate, or weak on coupling, real-time capability, scalability, extensibility, complexity, and security. These tables do the argument's work: they convert a set of system examples into a decision framework and ground the paper's later identification of pattern-specific open challenges.
What would settle it
Catalog every DBMS-LLM integration described in a broad corpus, such as all systems mentioned in recent published surveys, and check each against the five patterns; if a substantial number of working systems fit no pattern without forcing, or if two of the five patterns consistently appear together in practice, the taxonomy's claim to represent the landscape fails.
Extended reading notes
Core claim
The central claim is that DBMS-LLM integration in industrial and business settings is not a single technique but a landscape of five architectural patterns, each with a distinct structural relationship between the two systems. In DB-first integration the LLM is embedded inside the database as an interface, optimizer, UDF, or native operator; in LLM-first integration the database acts as a retrieval or cache backend for an LLM-driven application; middle-layer integration inserts an orchestrator between them; pipe-connected integration links them as independent services through data pipelines; and platform-based integration provides both as managed services on one cloud platform. The paper further claims that these patterns differ systematically along dimensions such as coupling degree, real-time performance, scalability, extensibility, complexity, and security, and that no single pattern dominates: the right choice depends on task requirements, functional needs, LLM dependencies, and deployment mode. A corollary of the taxonomy is that hybrid combinations are common and that the categories are explicitly non-orthogonal.
Load-bearing premise
The five-way taxonomy is assumed to be complete and representative, but the paper never derives it from an exhaustive survey or a formal design space; if a major integration mode is missing or the categories overlap too much to be useful, the main contribution collapses.
Editorial extensions
If this is right
- DB-first integration, with LLMs as operators or UDFs, can extend SQL to semantic and multimodal operations while keeping database governance and indexing benefits.
- LLM-first integration suits conversational and retrieval-augmented generation applications but sacrifices fine-grained query control and transactional guarantees.
- Middle-layer integration makes pipelines composable and modular at the cost of added complexity, latency, and debugging difficulty.
- Pipe-connected integration fits event-driven, high-throughput scenarios but accepts higher latency and harder failure recovery.
- Platform-based integration lowers setup cost and offers scalability but risks vendor lock-in and limited transparency.
Reading between the lines
- A testable extension implied by the paper's non-orthogonality admission is that hybrid patterns, not pure ones, may dominate real deployments, so future work could quantify how often production systems mix two or more of the five patterns.
- The taxonomy could be operationalized as a selection agent: given workload latency, data volume, security needs, and user expertise, it could recommend one pattern; the paper mentions this as future work but does not specify the decision procedure.
- The paper's cost-modeling challenge suggests a measurable benchmark: compare end-to-end query latency and accuracy of DB-first versus middle-layer integration on the same hybrid workload, something no standardized benchmark currently exists for.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys recent work on integrating Database Management Systems (DBMSs) and Large Language Models (LLMs) in industrial and business applications. Its main contribution is a taxonomy of five architectural patterns — DB-first, LLM-first, Middle-layer, Pipe-connected, and Platform-based — along with qualitative comparisons, selection guidance, and a discussion of open challenges. The paper makes no empirical claims; the value rests on the coherence and usefulness of the proposed taxonomy and on the accuracy of the survey coverage.
Significance. If the proposed five-way taxonomy were convincing, the paper would provide a useful organizing framework for a rapidly growing area and could help practitioners choose among integration strategies. The survey covers a broad set of recent systems and honestly acknowledges that the categories are non-orthogonal and that hybrid combinations are common. The comparison tables and the enumerated future challenges are useful entry points. However, the taxonomy is the central load-bearing claim, and as detailed in the major comments it currently mixes heterogeneous classification criteria and contains internal inconsistencies. The contribution is therefore promising but not yet demonstrated in the present form.
major comments (4)
- [Section 3.1, Figure 1, Table 2] The five categories are not derived systematically from the three dimensions stated in Section 3.1 (purposes, system layer, coupling degree). The categories mix at least four distinct criteria: coupling (DB-first vs. LLM-first), presence of an orchestration layer (Middle-layer), dataflow topology (Pipe-connected), and deployment environment (Platform-based). The manuscript itself concedes the categorization is non-orthogonal. As a result, the central claim that these are 'five representative architectural patterns' is not established; the paper currently presents overlapping perspectives rather than a partition of a design space. Please either derive the categories from a single coherent set of dimensions or explicitly reframe them as orthogonal concerns rather than mutually exclusive patterns.
- [Section 3.6 vs. Section 3.2] The Platform-based category is not an architectural pattern in the same sense as the other four; it denotes managed cloud offerings (e.g., Snowflake Cortex, Google Cloud, Oracle Cloud) that can host any of the other patterns. For example, an LLM-as-UDF system described in Section 3.2 would belong to both DB-first and Platform-based if it were deployed on a cloud platform such as Snowflake Cortex. This conflation weakens the taxonomy and the selection guidance in Section 3.7. Please clarify whether Platform-based is a deployment axis orthogonal to the other four and adjust the claims and Table 2 accordingly.
- [Section 3.7 vs. Table 3] The selection guidance contradicts the comparison table. Section 3.7 states that 'Real-time and low-latency tasks are better served by DB-first, pip-connected, or platform-based integrated solutions,' but Table 3 rates Pipe-connected as ≈ (Moderate) on Real-Time, and Table 2 lists 'High latency or eventual consistency' as a weakness of Pipe-connected. This inconsistency is directly relevant to the paper's practical guidance and should be reconciled.
- [Table 3] The qualitative Strong/Moderate/Weak scores in Table 3 are presented without a methodology or supporting citations. The 'Complexity' row is especially ambiguous: for most rows a Strong score is desirable, but for Complexity a Strong score would presumably be undesirable, yet the symbol legend defines ✓/≈/× only as Strong/Moderate/Weak. Since the paper's practical contribution includes choosing integration strategies (Section 3.7), these unsupported and sign-ambiguous scores weaken the advice and need justification or a clear convention.
minor comments (5)
- [Throughout references] Several citations use given names instead of surnames, e.g., 'Kaushikpresent et al. [53]' (should be Rajan et al.), 'Moreh et al. [48]' (should be Park et al.), 'Jinyang et al. [35]' (should be Li et al.), 'Zhaodonghui et al. [36]' (Li et al.), 'Wenbo et al. [60]' (Sun et al.), 'Alekh et al. [31]' (Jindal et al.), 'Xinyang et al. [79]' (Zhao et al.), 'Sumedh et al. [54]' (Rasal et al.), 'Jiayi et al. [72]' (Wang et al.), 'Kai et al. [71]' (Waehner et al.), and 'Simone et al. [46]' (Papicchio et al.). Please standardize to surname-based citations.
- [Equation (1)] Equation (1), 'LLMs + DBMSs → Dual Infrastructures of Enterprises,' is not a mathematical equation and adds no information beyond the prose; consider removing it or rewriting the statement in words.
- [Figure 1] Figure 1 is very dense, and some labels such as 'DB-side fusion' do not match the terminology used elsewhere ('DB-first'). The arrows for data flow versus control flow are hard to distinguish at print size. Please enlarge the figure, align labels with the text, and make the legend more readable.
- [Section 2] The related work section lists prior surveys but does not clearly differentiate the proposed five-category taxonomy from the three paradigms of [82] or the interaction paradigms of [34]. A short positioning paragraph or a comparison table would strengthen the novelty claim.
- [Section 3.6 and Section 3.7] There is a typo in Section 3.6: 'This create significant opportunities' should be 'This creates significant opportunities.' Also, 'pip-connected' appears in Section 3.7, which is inconsistent with 'Pipe-connected' used elsewhere.
Circularity Check
No significant circularity: the five-pattern taxonomy is a survey classification built from cited systems and qualitative trade-offs, not a derivation that reduces to its own inputs.
full rationale
This paper is a survey, not a derivation or prediction pipeline. Its central claim is a taxonomy of five DBMS-LLM integration strategies, introduced in Section 3.1 as follows: 'Based on these dimensions, we identify five major integration strategies, as illustrated in Figure 1.' The categories are then populated with examples from the surveyed literature, and the paper explicitly notes that 'our categorization is non-orthogonal' and that hybrid combinations occur. That is standard survey classification practice, not circular reasoning. No equation is derived from an input and then relabeled as a prediction; no fitted parameter is used as an output; and no uniqueness theorem or prior result is invoked to force the taxonomy. The only author self-citation is [75] in the Related Work section ('Yan et al. [75] focus specifically on DRL-based techniques for join order selection, a key challenge in query optimization'), which is background context and is not load-bearing for the five-pattern claim. Concerns about whether the five categories are complete, mutually consistent, or derived cleanly from the stated dimensions are validity and framing concerns, not circularity; they do not amount to a reduction of the paper's conclusions to its assumptions. Accordingly, no circular step can be exhibited, and the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper The five proposed categories (DB-first, LLM-first, middle-layer, pipe-connected, platform-based) are assumed to be representative and useful.
- domain assumption Qualitative ratings in Table 3 (Strong/Moderate/Weak) are assumed accurate.
- domain assumption The cited papers accurately represent the state of the art in DBMS-LLM integration.
Cite this review
Pith. "Pith review of DBMS-LLM Integration Strategies in Industrial and Business Applications: Current Status and Future Challenges." pith.science (2026). https://pith.science/paper/Y22VT35A
@misc{pith2026250719254,
author = {Pith},
title = {Pith review of: DBMS-LLM Integration Strategies in Industrial and Business Applications: Current Status and Future Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y22VT35A}},
note = {Machine review of arXiv:2507.19254}
}
read the original abstract
Modern enterprises are increasingly driven by the DATA+AI paradigm, in which Database Management Systems (DBMSs) and Large Language Models (LLMs) have become two foundational infrastructures powering a wide range of industrial and business applications, such as enterprise analytics, intelligent customer service, and data-driven decision-making. The efficient integration of DBMSs and LLMs within a unified system offers significant opportunities but also introduces new technical challenges. This paper surveys recent developments in DBMS+LLM integration and identifies key future challenges. Specifically, we categorize five representative architectural patterns based on their core design principles, strengths, and trade-offs. Based on this analysis, we further highlight several critical open challenges. We aim to provide a systematic understanding of the current integration landscape and to outline the unresolved issues that must be addressed to achieve scalable and efficient integration of traditional data management and advanced language reasoning in future intelligent applications.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 (2023). https://doi.org/10.48550/arXiv.2303.08774
-
[2]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al
-
[3]
Amazon Aurora. [n.d.]. https://aws.amazon.com/cn/rds/aurora/
-
[4]
Mouna Ammar, Christopher Rost, Riccardo Tommasini, Shubhangi Agarwal, Angela Bonifati, Petra Selmer, and Erhard Rahm. 2025. Towards Hybrid Graphs: Unifying Property Graphs and Time Series. In 28th International Conference on Extending Database Technology (EDBT) . 2483–2490. https://openproceedings. org/2025/conf/edbt/paper-183.pdf
2025
-
[5]
Anthropic: Claude. [n.d.]. https://www.anthropic.com/
-
[6]
Apache Flink. [n.d.]. https://flink.apache.org/
-
[7]
Apache Kafka. [n.d.]. https://kafka.apache.org/
-
[8]
Apache Spark. [n.d.]. https://spark.apache.org/
Show all 89 references
-
[9]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen Technical Report. arXiv preprint arXiv:2309.16609 (2023)
2023 arXiv
-
[10]
BigQuery. [n.d.]. https://cloud.google.com/bigquery
-
[11]
Asim Biswal, Liana Patel, Siddarth Jha, Amog Kamsetty, Shu Liu, Joseph E Gonzalez, Carlos Guestrin, and Matei Zaharia. 2025. Text2SQL is Not Enough: Unifying AI and Databases with TAG. In CIDR
2025
-
[12]
Qingpeng Cai, Can Cui, Yiyuan Xiong, Wei Wang, Zhongle Xie, and Meihui Zhang. 2022. A Survey on Deep Reinforcement Learning for Data Processing and Analytics. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4446–4465
2022
-
[13]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gau- rav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sut- ton, Sebastian Gehrmann, et al . 2023. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research 24, 240 (2023),...
2023
-
[14]
ClickHouse. [n.d.]. https://clickhouse.com/
-
[15]
Edgar F Codd. 1970. A Relational Model of Data Large Shared Data Banks. Commun. ACM 13, 6 (1970), 377–387
1970
-
[16]
Haowen Dong, Chao Zhang, Guoliang Li, and Huanchen Zhang. 2024. Cloud- Native Databases: A Survey. IEEE Transactions on Knowledge and Data Engineer- ing 36, 12 (2024), 7772–7791. https://doi.org/10.1109/TKDE.2024.3397508
2024
-
[17]
Anas Dorbani, Sunny Yasser, Jimmy Lin, and Amine Mhedhbi. 2025. Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB. arXiv preprint arXiv:2504.01157 (2025)
2025 arXiv
-
[18]
DuckDB. [n.d.]. https://duckdb.org/
-
[19]
Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its Nature, Scope, Limits, and Consequences. Minds and Machines 30, 4 (2020), 681–694
2020
-
[20]
Yannis Foufoulas and Alkis Simitsis. 2023. Efficient Execution of User-Defined Functions in SQL Queries. Proceedings of the VLDB Endowment 16, 12 (2023), 3874–3877
2023
-
[21]
Kai Franz, Samuel Arch, Denis Hirn, Torsten Grust, Todd C Mowry, and Andrew Pavlo. 2024. Dear User-Defined Functions, Inlining isn’t working out so great for us. Let’s try batching to make our relationship work. Sincerely, SQL. In Conference on Innovative Data Systems Research (CIDR)
2024
-
[22]
Parker Glenn, Parag Dakle, Liang Wang, and Preethi Raghavan. 2024. BlendSQL: A Scalable Dialect for Unifying Hybrid Question Answering in Relational Algebra. In Findings of the Association for Computational Linguistics ACL 2024 . 453–466
2024
-
[23]
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. 2024. ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools. arXiv preprint arXiv:2406.12793 (2024)
2024 arXiv
-
[24]
Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, and Yu Su. 2024. Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments. arXiv preprint arXiv:2402.14672 (2024)
2024 arXiv
-
[25]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al . 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948 (2025)
2025 arXiv
-
[26]
Yikun Han, Chunjiang Liu, and Pengfei Wang. 2023. A Comprehensive Survey on Vector Database: Storage and Retrieval Technique, Challenge. arXiv preprint arXiv:2310.11703 (2023)
2023
-
[27]
Mohamed S Hassan, Tatiana Kuznetsova, Hyun Chai Jeong, Walid G Aref, and Mohammad Sadoghi. 2018. GRFusion: Graphs as First-Class Citizens in Main- Memory Relational Database Systems. In Proceedings of the 2018 International Conference on Management of Data . 1789–1792. https:/...
2018 doi
-
[28]
Zijin Hong, Zheng Yuan, Qinggang Zhang, Hao Chen, Junnan Dong, Feiran Huang, and Xiao Huang. 2024. Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL. arXiv preprint arXiv:2406.08426 (2024)
2024
- [29]
-
[30]
IBM Db2. [n.d.]. https://www.ibm.com/products/db2
-
[31]
Alekh Jindal, Shi Qiao, Sathwik Madhula, Kanupriya Raheja, and Sandhya Jain
- [32]
- [33]
- [34]
-
[35]
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al. 2023. Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to- SQLs. Advances in Neural Information Processing Sy...
2023
-
[36]
Zhaodonghui Li, Haitao Yuan, Huiming Wang, Gao Cong, and Lidong Bing. 2024. LLM-R2: A Large Language Model Enhanced Rule-Based Rewrite System for Boosting Query Efficiency. Proceedings of the VLDB Endowment 18, 1 (2024), 53–65
2024
-
[37]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. DeepSeek-V3 Technical Report. arXiv preprint arXiv:2412.19437 (2024)
2024 arXiv
-
[38]
Microsoft SQL Server. [n.d.]. https://www.microsoft.com/en-us/sql-server
-
[39]
Mistral. [n.d.]. https://mistral.ai/
-
[40]
MySQL. [n.d.]. https://www.mysql.com/
-
[41]
Neo4j. [n.d.]. https://neo4j.com/
-
[42]
OpenAI: CLIP. [n.d.]. https://openai.com/index/clip/
-
[43]
Oracle Database. [n.d.]. https://www.oracle.com/database/
-
[44]
James Jie Pan, Jianguo Wang, and Guoliang Li. 2024. Survey of Vector Database Management Systems. The VLDB Journal 33, 5 (2024), 1591–1615
2024
-
[45]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu
- [46]
-
[47]
Paolo Papotti. 2024. Querying Structured and Unstructured Data: LLM-first or DB-first? Keynote Talk of 29th International Conference on Database Systems for Advanced Applications (DASFAA)
2024
-
[48]
IEEE Transactions on Knowledge and Data Engineering 36, 7 (2024), 3580–3599
Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering 36, 7 (2024), 3580–3599. https://doi.org/10.1109/TKDE.2024.3352100
2024
-
[49]
Linnea Passing, Manuel Then, Nina C Hubig, Harald Lang, Michael Schreier, Stephan Günnemann, Alfons Kemper, and Thomas Neumann. 2017. SQL-and Operator-centric Data Analytics in Relational Main-Memory Databases. In 20th International Conference on Extending Database Technology ...
2017
-
[50]
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding Multimodal Large Language Models to the World. arXiv preprint arXiv:2306.14824 (2023)
2023 arXiv
-
[51]
Hyunbyung Park, Sukyung Lee, Gyoungjin Gim, Yungi Kim, Dahyun Kim, and Chanjun Park. 2024. Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models. arXiv preprint arXiv:2403.19340 (2024)
2024 arXiv
-
[52]
PostgreSQL. [n.d.]. https://www.postgresql.org/
-
[53]
Kaushik Rajan, Aseem Rastogi, Akash Lal, Sampath Rajendra, Krithika Subrama- nian, and Krut Patel. 2024. Welding Natural Language Queries to Analytics IRs with LLMs. In CIDR
2024
-
[54]
Maximilian Plazotta and Meike Klettke. 2024. Data Architectures in Cloud Environments. Datenbank-Spektrum (2024), 243–247. https://doi.org/10.1007/ s13222-024-00490-5
2024
-
[55]
Mohammed Saeed, Nicola De Cao, and Paolo Papotti. 2023. A DB-First Approach to Query Factual Information in LLMs. In NeurIPS 2023 Second Table Representa- tion Learning Workshop. https://openreview.net/forum?id=R8VFPAfOcN
2023
-
[56]
Mohammed Saeed, Nicola De Cao, and Paolo Papotti. 2024. Querying Large Lan- guage Models with SQL. In 27th International Conference on Extending Database Technology (EDBT). 365–372. https://openproceedings.org/2024/conf/edbt/paper- 61.pdf
2024
-
[57]
Sumedh Rasal. 2024. A Multi-LLM Orchestration Engine for Personalized, Context-Rich Assistance. arXiv preprint arXiv:2410.10039 (2024)
2024 arXiv
-
[58]
Snowflake. [n.d.]. https://www.snowflake.com/en/
-
[59]
Michael Stonebraker and Andrew Pavlo. 2024. What Goes Around Comes Around... And Around... ACM Sigmod Record 53, 2 (2024), 21–37
2024
-
[60]
Moritz Sichert and Thomas Neumann. 2022. User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern Databases. Proceedings of the VLDB Endowment 15, 5 (2022), 1119–1131. https://doi.org/10.14778/3510397.3510408
2022
-
[61]
Zhaoyan Sun, Xuanhe Zhou, and Guoliang Li. 2024. R-Bot: An LLM-based Query Rewrite System. arXiv preprint arXiv:2412.01661 (2024)
2024 arXiv
-
[62]
Jie Tan, Kangfei Zhao, Rui Li, Jeffrey Xu Yu, Chengzhi Piao, Hong Cheng, Helen Meng, Deli Zhao, and Yu Rong. 2025. Can Large Language Models Be Query Optimizer for Relational Databases? arXiv preprint arXiv:2502.05562 (2025)
2025 arXiv
-
[63]
Wenbo Sun, Ziyu Li, Vaishnav Srinidhi, and Rihan Hai. 2025. Database is All You Need: Serving LLMs with Relational Queries. In International Conference on Extending Database Technology (EDBT) . 1118–1121. https://openproceedings. org/2025/conf/edbt/paper-326.pdf
2025
-
[64]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530 (2024)
2024 arXiv
- [65]
-
[66]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al
-
[67]
Matthias Urban and Carsten Binnig. 2024. Demonstrating CAESURA: Language Models as Multi-Modal Query Planners. In Companion of the 2024 International Conference on Management of Data . 472–475. https://doi.org/10.1145/3626246. 3654732
2024 doi
- [68]
-
[69]
Matthias Urban and Carsten Binnig. 2024. ELEET: Efficient Learned Query Execution over Text and Tables. Proceedings of the VLDB Endowment 17, 13 (2024), 4867–4880. https://doi.org/10.14778/3704965.3704989
2024
-
[70]
Matthias Urban and Carsten Binnig. 2024. CAESURA: Language Models as Multi-Modal Query Planners. In Conference on Innovative Data Systems Research (CIDR). https://vldb.org/cidrdb/papers/2024/p14-urban.pdf
2024
-
[71]
Kai Waehner. 2025. The Ultimate Data Streaming Guide: Concepts, Use Cases, Industry Stories. Confluent
2025
-
[72]
Jiayi Wang and Guoliang Li. 2025. AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries. In Conference on Innovative Data Systems Research (CIDR). https://vldb.org/cidrdb/papers/2025/p32-wang.pdf
2025
-
[73]
Tianshu Wang, Xiaoyang Chen, Hongyu Lin, Xianpei Han, Le Sun, Hao Wang, and Zhenyu Zeng. 2025. DBCopilot: Natural Language Querying over Massive Databases via Schema Routing. In 28th International Conference on Extending Database Technology (EDBT). 707–721. https://openproceed...
2025
-
[74]
Nikita Vasilenko, Alexander Demin, and Vladimir Boorlakov. 2025. Training- Free Query Optimization via LLM-Based Plan Similarity. arXiv preprint arXiv:2506.05853 (2025)
2025 arXiv
-
[75]
Zhengtong Yan, Valter Uotila, and Jiaheng Lu. 2023. Join Order Selection with Deep Reinforcement Learning: Fundamentals, Techniques, and Challenges. Pro- ceedings of the VLDB Endowment 16, 12 (2023), 3882–3885
2023
-
[76]
Zhiming Yao, Haoyang Li, Jing Zhang, Cuiping Li, and Hong Chen. 2025. A Query Optimization Method Utilizing Large Language Models. arXiv preprint arXiv:2503.06902 (2025)
2025 arXiv
-
[77]
Fuheng Zhao, Divyakant Agrawal, and Amr El Abbadi. 2025. Hybrid Querying Over Relational Databases and Large Language Models. In Conference on Inno- vative Data Systems Research (CIDR) . https://vldb.org/cidrdb/papers/2025/p10- zhao.pdf
2025
- [78]
-
[79]
Xinyang Zhao, Xuanhe Zhou, and Guoliang Li. 2024. Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs. Proceedings of the VLDB Endowment 17, 12 (2024), 4481–4484. https://doi.org/10.14778/3685800. 3685905
2024 doi
-
[80]
Jiehan Zhou, Yang Cao, Quanbo Lu, Weishan Zhang, Xin Liu, and Weijian Ni
-
[81]
Jiehan Zhou, Yang Cao, Quanbo Lu, Yan Zhang, Cong Liu, Shouhua Zhang, and Junsuo Qu. 2024. Industrial Large Model: A Survey. InMATEC Web of Conferences, Vol. 401. EDP Sciences, 10009. https://doi.org/10.1051/matecconf/202440110009
2024
- [82]
-
[83]
Xuanhe Zhou, Chengliang Chai, Guoliang Li, and Ji Sun. 2020. Database Meets Artificial Intelligence: A Survey. IEEE Transactions on Knowledge and Data Engineering 34, 3 (2020), 1096–1116
2020
-
[84]
Xuanhe Zhou, Xinyang Zhao, and Guoliang Li. 2024. LLM-Enhanced Data Management. arXiv preprint arXiv:2402.02643 (2024)
2024 arXiv
-
[85]
In 2024 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE)
Industrial Large Model: Toward A Generative AI for Industry. In 2024 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE) . IEEE, 80–81
2024
-
[87]
Lixi Zhou, Qi Lin, Kanchan Chowdhury, Saif Masood, Alexandre Eichenberger, Hong Min, Alexander Sim, Jie Wang, Yida Wang, Kesheng Wu, et al. 2024. Serving Deep Learning Models from Relational Databases. In27th International Conference on Extending Database Technology (EDBT) . 7...
2024
-
[2022]
Advances in neural information processing systems 35 (2022), 23716–23736
Flamingo: a Visual Language Model for Few-Shot Learning. Advances in neural information processing systems 35 (2022), 23716–23736
2022
-
[2023]
arXiv preprint arXiv:2312.11805 (2023)
Gemini: A Family of Highly Capable Multimodal Models. arXiv preprint arXiv:2312.11805 (2023)
2023 arXiv
-
[2024]
In Conference on Innova- tive Data Systems Research (CIDR)
Turning Databases Into Generative AI Machines. In Conference on Innova- tive Data Systems Research (CIDR)
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.