{"id":"dc91cf6b-fa2c-4d4b-ac54-77d12de4cd60","arxiv_id":"2505.13648","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A structured review of 5,300 conceptual modeling papers finds the field has shifted from data modeling toward process modeling, with gaps in AI, blockchain, and social media research.","lead":"This paper reviews five decades of research on conceptual modeling by analyzing over 5,300 papers from 35 journals and conferences. It shows how modeling topics have shifted over time and suggests where the field should go next in a digital world.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The trend claims rest on an unreconciled corpus subset: abstract/Table 2 say 5,303 papers, while Table 4's topic analysis covers 4,345 considered and 3,173 full-text papers, so the process-vs-data and AI/blockchain gap conclusions may be artifacts of which papers actually entered the model.","rationale":"The reader's weakest assumption was corpus representativeness. I agree that this is the load-bearing premise, and I sharpen it to a specific, internally visible defect: the claimed 5,303-paper corpus does not match the 4,345/3,173 totals used for the actual topic analysis, and the missing segment is not described. This is a stronger and more concrete version of the same concern. The paper is otherwise valuable: the qualitative synthesis is plausible, the mixed-method design is reasonable, and the CAiSE/EMISAJ comparison is a good internal check. But because every quantitative conclusion in Section 3 flows from yearly topic models, an unexplained one-fifth reduction in the analyzed corpus, with no source-year detail, leaves the central trend claims unverified. The requested crosswalk and robustness re-run would settle whether the concern is real or benign; if the trend holds under full-source accounting, the paper's claims are much stronger. I do not see grounds for rejection, and the reader's CONDITIONAL verdict already captures the appropriate stance, so no change is needed.","tokens_in":34819,"tokens_out":6180,"duration_ms":57950,"concrete_test":"Produce a source-by-year crosswalk of the corpus: for each of the 35 sources and each year, report the number of papers retrieved, the number manually screened as relevant, and the number with full text, so that Table 4's totals (4,345 considered; 3,173 full text) reconcile with Table 2's 5,303. Then re-run the Section 3 LDA and term-prevalence analyses on the reported subset and compare with a version that restricts to the 2005-2020 window where full-text coverage is complete; if the relative prevalence of process-related versus data-related topics or the near-zero AI/blockchain shares shifts materially, the headline claims are artifacts of unstated exclusions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claims—a shift from data-oriented to process-oriented modeling, and underexplored AI/blockchain/social-media topics—depend entirely on the corpus fed into the LDA and term-prevalence analyses in Section 3. That corpus is not transparently delimited. Section 2.2.3 states that the 'final number of papers analyzed was 5,303 across 35 journals and conferences,' and Table 2 sums to exactly 5,303. Yet Table 4, which reports the yearly buckets used for topic analysis, sums to 4,345 'articles considered' and only 3,173 'articles with full text.' No reconciliation is provided: we are not told which 958 papers are absent from Table 4, nor which 1,172 papers lacked full text. The missing ER full texts before 2005, flagged in the Table 2 footnote, are a concrete instance: Table 2 lists 2,062 ER papers, but the 1976-2004 row of Table 4 lists only 801 considered and 639 full-text papers. If the early period is built on a small, non-representative slice that excludes the field's flagship data-modeling venue, the headline 'data-oriented then, process-oriented now' contrast may be an artifact of availability, not a real trend. Likewise, the near-zero prevalence of AI, blockchain, and social-media terms could reflect keyword and venue choices rather than genuine gaps: the 35 sources are largely IS/SE/database venues, and communities that publish on AI or blockchain often use different vocabulary (e.g., 'schema,' 'ontology,' 'model card') and different outlets. The internal-consistency check in Table 8 (CAiSE/EMISAJ similarity to the rest of the corpus) does not answer this concern, because CAiSE is already part of the main corpus and EMISAJ is another outlet from the same modeling community; it validates internal overlap, not external representativeness. This is an internal consistency problem, not a disagreement with consensus, and it is directly checkable from the data the authors should have kept.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a structured literature review of conceptual modeling research from 1976 to 2022. The authors assembled a corpus of 5,303 papers from 35 journals and conferences, applied manual screening with an explicit inclusion protocol, ran yearly LDA topic models and a Doc2Vec language model on full texts, and supplemented the quantitative results with qualitative interpretation. The central claims are that conceptual modeling has shifted from data-oriented to process-oriented modeling over the past fifteen years; that core themes such as ER modeling, UML, BPMN, and ontology remain stable; and that technologies such as AI, blockchain, social media, and goal modeling are underexplored relative to their societal importance. The paper closes with a research agenda organized around foundations, non-traditional settings, new frameworks, modeling processes, a broadened user base, and grammar–script relationships.","tokens_in":35190,"tokens_out":5213,"duration_ms":46966,"significance":"The paper offers a large and carefully motivated corpus, and if the empirical claims hold, it would be a useful reference point for the conceptual modeling community. The explicit inclusion protocol, manual screening by multiple coders, comparison with Härer and Fill, and corpora-similarity check using sentence embeddings are genuine strengths that go beyond many prior reviews. The claimed shift from data-oriented to process-oriented modeling and the identification of underexplored areas are consequential for future research agendas. However, the paper's transparency is currently insufficient to support those claims at the level of confidence implied by the abstract, because the corpus totals are unreconciled and the modeling pipeline is not reproducible from the information given.","major_comments":[{"comment":"The paper does not reconcile the corpus totals. Section 2.2.3 and Table 2 state that the final corpus contains 5,303 papers, but Table 4 reports only 4,345 'articles considered' and 3,173 'articles with full text' for the yearly analysis. Since the LDA and term-prevalence analyses in Section 3 are run on the full-text subset, the abstract's 'over 5,300 papers' overstates the evidence base for the central trend claims. The authors should provide a per-source and per-year reconciliation of the 958 missing papers, state explicitly which analyses use which subsets, and justify the decision to focus the statistical analysis on 2005–2020 in Section 3 while Table 4 also reports earlier years with lower full-text availability.","section":"Section 2.2.3, Table 2 vs. Table 4"},{"comment":"The early-period comparison is not supported by the available full texts. Table 2 lists 2,062 ER conference papers, but the footnote states that full texts before 2005 were not available; Table 4 shows the entire 1976–2004 bucket as 801 considered and 639 full-text papers across all venues. This means the claim that early conceptual modeling was predominantly data-oriented rests on a possibly non-representative subset, because the flagship venue's early output is missing. The authors should quantify the missing coverage, report topic models separately for ER and non-ER sources, or provide a sensitivity analysis that omits the pre-2005 period from the data-to-process narrative.","section":"Section 3.1, Table 2 footnote, Table 4"},{"comment":"The LDA pipeline is not reproducible and its outputs are not validated. The number of topics is selected by highest coherence for each year with a cap at 15, but the paper reports coherence only for year 2020 in Figure 2; it does not report the selected k for each year or the coherence values for all years, nor does it report LDA stability across random seeds or hyperparameter settings. Doc2Vec is described as built on 2,555 full papers (Table 3, Step f), which contradicts the 3,173 total in Table 4. The authors should release the code, the corpus identifiers, and the topic-term distributions for all years, and provide stability or robustness checks; without these, the interpretive topic labels in Sections 3.1–3.2 cannot be distinguished from researcher judgment.","section":"Section 2.3, Table 3"},{"comment":"The gap conclusions for AI, blockchain, and social media are stated as findings about the field, but the evidence only supports absence from the selected venues and keyword protocol. The sources are dominated by IS/SE/database outlets, and the inclusion keywords listed in Table 1 are terms like 'conceptual model', 'entity-relationship', and 'process'; they do not include terms such as 'model card', 'schema', 'ontology', or 'data model' that AI and blockchain communities also use. The paper should explicitly qualify all gap statements as corpus-relative and, ideally, validate them by running the same topic model on a supplementary sample from AI/blockchain venues or by using broader search terms.","section":"Sections 3.2.2, 3.4, 5"}],"minor_comments":[{"comment":"Table 3 expands LDA as 'linear discriminant analysis,' but Section 2.3 and the rest of the paper use LDA for Latent Dirichlet Allocation; please correct the expansion or distinguish the two methods.","section":"Table 3, Step b"},{"comment":"The meaning of an X in Table 6 is not defined; the paper does not state a frequency threshold or normalization criterion that qualifies a term for inclusion, so the table is not reproducible without additional detail.","section":"Table 6"},{"comment":"The column 'Frequency (this work)' repeats the same value across multiple topics (e.g., model = 61 appears in every topic row), which makes it unclear whether these are corpus-wide frequencies or per-topic frequencies; clarify how the frequencies were computed.","section":"Table 10"},{"comment":"The sentenceTransformer similarity analysis reports global overlap scores but does not state what text was embedded (topic labels, topic-term distributions, or full documents) or why 0.784 is interpreted as 'substantial similarity'; include the comparison details.","section":"Section 3.3, Table 8"},{"comment":"Several passages contain encoding artifacts (e.g., the 'we used full-text search as opposed to' sentence in Section 2.2.3 and the 's' fragments in Section 3.3); these should be cleaned before publication.","section":"Section 2.2.3 and Section 3.3"},{"comment":"Table 4 includes 2021 and 2022 rows with only CAiSE and EMISAJ papers (footnotes 3 and 4), which is inconsistent with the abstract's claim of coverage to 'the present' and with the stated focus on 2005–2020; state how these years are used or remove them.","section":"Table 4 and Section 3"},{"comment":"The text refers to Appendix A as containing the entire list of papers considered, but the list is not included in the submitted manuscript; if it is in an online supplement, say so explicitly and provide access.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"Given the paper's reliance on a corpus that is not publicly listed and whose totals are inconsistent, I recommend that the editor require the authors to deposit the full paper list, per-source and per-year counts, and analysis code as a condition of revision. The conclusions about field-level gaps (AI, blockchain, social media) should be scaled back until validated on broader sources or clearly re-framed as corpus-relative statements."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. This is the broadest map of conceptual modeling research I've seen, and the central narrative—a shift from data-oriented to process-oriented modeling—is plausible. But the corpus arithmetic doesn't reconcile, no data are shipped, and the quantitative conclusions are therefore weaker than the prose suggests.\n\nWhat's genuinely new: the corpus size (5,303 papers from 35 venues), the manual pre-screening layered onto LDA, the Doc2Vec augmentation for semantic similarity, and the separate CAiSE/EMISAJ comparison. The topic tables by year are a useful reference. The paper also positions itself well against Härer and Fill and Recker et al., which helps the reader see what's incremental and what's not.\n\nThe soft spots are real but fixable. Table 2 sums to 5,303, while Table 4 sums to 4,345 considered and 3,173 full-text papers; the gap is never explained. Which 958 papers dropped? Which 1,172 lacked full text? The ER full-text gap before 2005 is concrete: Table 2 lists 2,062 ER papers, but the 1976-2004 row of Table 4 has only 801 and 639. Since ER is the flagship data-modeling venue, the early 'data-oriented' slice may be a non-representative minority, which puts the headline trend at risk. The CAiSE/EMISAJ similarity check (Table 8) doesn't address this; it shows internal overlap, not external representativeness. No code, no data, no LDA stability metrics, and topic labels are interpretive. These are moderate issues for a survey, not fatal ones—the central claim is still the most reasonable reading of the evidence available, but the specific percentages should be read as provisional.\n\nThis paper is for IS and database researchers who want a comprehensive starting point on the field's trajectory. It is not a new mechanism or a falsifiable prediction, so judge it on transparency and usefulness. It deserves a serious referee; a good referee would ask for a reconciled data appendix and a breakdown of the full-text availability gap before the strong quantitative statements go to press.\n\nRecommendation: engage with it, cite it with caveats, and push for the data release in review.","headline":"A useful large-scale map of conceptual modeling research, but the headline trend rests on an unreconciled corpus subset and lacks shipped data; the core narrative is plausible, not confirmed.","tokens_in":35790,"tokens_out":3294,"would_cite":true,"duration_ms":30723,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 5,303-paper survey shows conceptual modeling shifting from data to process models.","keywords":["conceptual modeling","structured literature review","topic modeling","Latent Dirichlet Allocation","process modeling","data modeling","emerging technologies","information systems"],"falsifier":"Rebuild the corpus from a venue mix that excludes the most process-oriented conferences, then re-run the same per-year LDA; if process topics no longer dominate the post-2005 clusters, the central shift claim fails, and if AI-related papers appear in large numbers in machine-learning venues under the same keyword rules, the 'AI is underexplored' gap is a sampling artifact.","tokens_in":34634,"feed_emoji":"📊","tokens_out":5576,"duration_ms":50311,"temperature":0.7,"pith_summary":"This paper surveys more than 5,300 papers across 35 journals and conferences, spanning 1976 to 2022, to establish how conceptual modeling research has evolved over five decades. Its central claim is that conceptual modeling remains anchored in a core of entities, relationships, and processes, even as the field's center of gravity migrated from data-oriented modeling (entity-relationship diagrams, UML class models) to process-oriented modeling (BPMN, process mining) over the past fifteen years. The survey also establishes that major emergent technologies — artificial intelligence, social media, blockchain, and goal- and value-oriented modeling — attract far less conceptual modeling research than their digital prominence would suggest. A sympathetic reader would see this as evidence that the discipline is mature but in need of new frameworks and broader participation.","feed_headline":"Survey finds conceptual modeling shifted from data to process models","feed_subtitle":"A 5,303-paper analysis shows AI, blockchain, and social media under-modeled despite the digital turn.","key_machinery":"The load-bearing mechanism is a structured corpus of 5,303 full-text papers collected from 35 journals and conferences, manually screened by six coders, then analyzed per year. Latent Dirichlet Allocation (a probabilistic method that represents each document as a mixture of topics, each topic as a distribution over words) produces the year-level topic clusters; coherence scores select the number of topics, and a Doc2Vec language model learns semantic similarity between terms such as BPMN and BPEL. The per-year topic clusters are what support the claims about the data-to-process shift and about which topics are absent.","core_discovery":"The paper's discovery, on its own terms, is a documented shift in what conceptual modeling research actually studies. Using full-text Latent Dirichlet Allocation topic models run year by year, together with a Doc2Vec language model trained on the corpus, the authors show that until about 2005 the dominant topics were data-oriented concerns — entity-relationship modeling, relational database design, and knowledge modeling — whereas from 2005 to 2020 roughly half of the topic clusters concern processes: business process design, process mining, and process analysis. The same analysis shows a persistent common core of modeling languages, ontologies, and methodological research, which the paper reads as the field's continuity. It then argues that goal, intention, and contingency modeling, along with applications to AI, social media, blockchain, and large language models, remain underexplored, and that the field must broaden its assumptions about who models and what modeling is for.","pith_inferences":["If the shift is real, then conceptual modeling's future value may lie in process and organizational contexts, while data representation is increasingly delegated to machine-learned or schema-less infrastructure.","The corpus's heavy weight of process-oriented venues (the ER conference, CAiSE, EMMSAD) could overstate the decline of data modeling; a corpus balanced with data-management venues might show a slower, more regional shift.","The Doc2Vec model the authors built is a reusable asset: tracking when terms such as 'large language model' or 'blockchain' begin to cluster with core modeling constructs would give a quantitative early signal of conceptual modeling absorbing a new technology.","A direct test of the survey's forward-looking claim is whether the 2023–2030 literature shows AI and goal modeling rising from their current small share."],"forward_implications":["Process modeling, especially BPMN and process mining, will likely remain the most active conceptual modeling research area for the near term.","Data-oriented conceptual modeling (ER, UML class diagrams) will continue to lose relative share as relational databases give ground to NoSQL and data lakes, unless new notations emerge.","Goal-, intention-, and value-oriented modeling (e.g., the i* tradition) remains a persistently under-filled niche despite decades of calls for more work.","Artificial intelligence, social media, blockchain, and large language models are open frontiers where conceptual modeling frameworks tailored to those settings are needed.","Conceptual modeling research should broaden from professional analysts to citizen modelers, with instance-based and narrative representations complementing abstraction-based diagrams."],"supporting_citations":[{"why":"Supplies the entity-relationship model, the foundational construct behind the early data-oriented era the survey documents.","marker":"[38]"},{"why":"The prior bibliometric analysis of conceptual modeling publications whose topic and term frequencies the survey compares against to validate its results.","marker":"[95]"},{"why":"The survey of Latent Dirichlet Allocation that grounds the topic-modeling method used to produce the yearly clusters.","marker":"[104]"},{"why":"The semantic data models survey that anchors the early data-modeling topic in the 1980s period.","marker":"[177]"},{"why":"The framework describing conceptual modeling's shift from representation to mediation in a digital world, which this survey's assumptions and future directions engage.","marker":"[184]"},{"why":"The classic conceptual modeling research agenda that the paper revisits and updates with corpus-level evidence.","marker":"[236]"}],"fun_headline_variants":["Modeling research shifts from data to process","5,300 papers chart modeling's data-to-process shift","From data to process: 50 years of modeling research","Process modeling rises as data focus fades","Survey finds modeling research pivot to processes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole trend story rests on the assumption that the 35 selected venues together with the keyword protocol and manual screening fairly represent what conceptual modeling research actually is, so if the venue list skews toward process-oriented outlets, the data-to-process shift and the AI/blockchain gaps are artifacts of source selection rather than field-wide facts.","fun_headline_variants_meta":{"raw":{"variants":["Modeling research shifts from data to process","5,300 papers chart modeling's data-to-process shift","From data to process: 50 years of modeling research","Process modeling rises as data focus fades","Survey finds modeling research pivot to processes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002077,"raw_usage":{"total_tokens":8035,"prompt_tokens":861,"completion_tokens":7174,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":7103}},"tokens_in":477,"tokens_out":7174,"duration_ms":42696,"temperature":1.0,"reasoning_tokens":7103,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:11:45.687949+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rebuild the corpus from a venue mix that excludes the most process-oriented conferences, then re-run the same per-year LDA; if process topics no longer dominate the post-2005 clusters, the central shift claim fails, and if AI-related papers appear in large numbers in machine-learning venues under the same keyword rules, the 'AI is underexplored' gap is a sampling artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The semantic data models survey that anchors the early data-modeling topic in the 1980s period."},{"cited_title":"MIS Quarterly, 2021","cited_arxiv_id":null,"evidence_quote":"The framework describing conceptual modeling's shift from representation to mediation in a digital world, which this survey's assumptions and future directions engage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The classic conceptual modeling research agenda that the paper revisits and updates with corpus-level evidence."}],"review_version":1}