{"id":"d4cf00e2-fa12-4c19-8a09-176454a85e01","arxiv_id":"2411.13722","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An SLR of 16 industrial reports shows TLA+ is mostly applied in cloud settings during early design, reports benefits in bug-finding and design quality, and notes steep learning curves and abstraction choices as key barriers.","lead":"A systematic review of 16 high-affinity reports finds that TLA+, a formal specification language, is used mainly in cloud systems, mostly during early design and debugging, and that companies report catching subtle bugs. The review maps the benefits and challenges of TLA+ in industry over 2013 to 2023, giving practitioners and researchers a curated starting point.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Concern: the RQ1 trend claim rests on the 261 Google Scholar papers without a baseline; the 16-paper industry corpus is too small for the 'surge since 2015' conclusion.","rationale":"The reader's weakest_assumption identifies essentially the same problem: the Google Scholar corpus of 261 deduplicated papers is the load-bearing basis for the RQ1 trend, and it includes irrelevant or nonsensical papers. My concern adds two specifics that strengthen the case. First, the query has no baseline or control condition. In a decade-long SLR of a topic like TLA+, the raw number of papers matching a broad query is likely to grow because Google Scholar indexes more venues, more preprints, and more grey literature over time, independent of industry adoption. A positive slope on such a count is therefore not evidence of an industry-specific trend unless compared to a control. Second, the paper's own data (Figure 1) show that the 16 high-affinity papers are irregularly distributed across years, so an 'upward trend' from that set would be fragile; the authors instead regress the 261-paper set and then generalize to industry. The paper is transparent about these limits in the Threats to Validity section (the 261 set 'may include irrelevant or nonsensical papers, which potentially skews our trend analysis'), which is a stated limitation in-scope and should weigh on the verdict. The strongest parts of the paper, its RQ2-RQ4 synthesis of the 16 curated papers and its published dataset, remain useful. The correct action is to keep the verdict CONDITIONAL, requiring authors to either add the baseline analysis or explicitly reframe RQ1 and the 'surged' conclusion as statements about Google Scholar publication volume, not industrial usage. I cannot accept the claim as stated, but the transparent methodology and honest limitation section make REJECT too strong. The key evidence needed is a simple re-analysis, not a new empirical study.","tokens_in":8041,"tokens_out":2037,"duration_ms":16590,"concrete_test":"Re-run the trend analysis with a control corpus. Take the same Google Scholar query without the industry terms from Listing 1 (or with 'industry' terms replaced by neutral terms), filter the same way, and compare the slope of the 261-paper trend to the slope of the control trend over 2013-2023. If the control slope is also significantly positive (p<0.05) and of comparable magnitude (within 20%), the 261-paper trend is not evidence of increasing industrial TLA+ usage and the RQ1 claim and Section 5 'surged' conclusion should be reframed as general publication-volume growth. Also run the regression on only the 16 high-affinity papers per year; if the small-sample regression shows a non-significant slope, the 'surge since 2015' claim lacks direct support.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in Section 3.1 ('significant upward trend in TLA+ papers in general, which is likely to continue due to success stories recommending its adoption in industry') and the Section 5 conclusion ('TLA+ usage has surged since 2015') are the load-bearing results. The trend regression is run on the 261 deduplicated papers that matched an optimistic Google Scholar query (Listing 1), which the authors themselves warn 'may include irrelevant or nonsensical papers.' The search string is not a measure of industrial TLA+ usage: it is a compound query over general terms like 'application', 'practice', 'industry', 'cloud', etc., so the 261-paper volume reflects research publication volume about TLA+ generally, not industry adoption. Without a baseline comparison (e.g., TLA+ papers without the industry terms, or PlusCal-only papers), the regression's upward slope is uninterpretable as evidence about industrial practice. The 16 high-affinity papers are too few for a per-year trend, most appear after 2018, and year-by-year counts vary widely (e.g., the 2020 spike and the 2013 drop). The conclusion 'surged since 2015' is not supported by any statistical analysis of the 16 industrial papers; it is an informal reading of the same 261-paper Google Scholar trend. Thus the paper's primary RQ1 answer and the broad 'surge' conclusion conflate publication volume with industrial usage and lack the baseline needed to support the claimed increasing industrial adoption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a systematic literature review (SLR) of industrial TLA+ practice from 2013 to 2023. It defines four research questions on trend, characterization, benefits, and challenges of TLA+ in industry. The authors use a Google Scholar search with a compound query (Listing 1), reduce 290 initial hits to 16 'high-affinity' papers via deduplication, exclusion, and an affinity filter, and then analyze those 16 papers. The paper claims a statistically significant upward trend in TLA+ publications (based on 261 deduplicated hits, S=3.67, p<0.001), characterizes industrial use as predominantly cloud-oriented and concentrated in early design/debugging, and reports that TLA+ delivers bug-finding and design-understanding benefits while facing steep learning curve and abstraction challenges. The filtered dataset is published on Zenodo.","tokens_in":8304,"tokens_out":6391,"duration_ms":55187,"significance":"If the trend claim were supported, the paper would be a valuable empirical contribution on formal-methods adoption in industry. The paper's strengths include a transparent and reproducible search string, a published dataset, independent dual coding with author adjudication, and a clear qualitative synthesis of 16 industrial case studies. The RQ2-RQ4 findings—cloud dominance, early-design usage, reported bug-finding benefits, and challenge themes—are useful for both practitioners and researchers. The main significance hinges on the RQ1 trend estimate, which currently conflates publication volume with industrial usage and lacks the data needed to support the 'surge' conclusion.","major_comments":[{"comment":"The central RQ1 answer—'There is a significant upward trend in TLA+ papers in general, which is likely to continue due to success stories'—and the concluding 'surge since 2015' are based on a linear regression of the 261 deduplicated Google Scholar hits, not on the 16 high-affinity papers that actually concern industrial use. The search string in Listing 1 is a broad compound query over general terms ('application', 'practice', 'insight', 'usage') and industry names, so the 261-paper series is a measure of publication volume matching those terms, not of industrial TLA+ adoption. The authors themselves state in Section 4 that this set 'may include irrelevant or nonsensical papers, which potentially skews our trend analysis.' No baseline comparison (e.g., TLA+ papers without industry terms, or another formal method) is provided to allow the upward slope to be attributed to industry practice. This is a load-bearing flaw in the paper's main contribution.","section":"Section 3.1 and Section 5"},{"comment":"The regression reporting is incomplete and the extrapolation is unsupported. The paper reports only the slope (S=3.67) and p-value; it gives no confidence intervals, R², model diagnostics, or sensitivity analysis with respect to the query wording or time range. More importantly, no trend test is run on the 16 high-affinity papers, and those counts are too sparse and variable (e.g., the 2020 peak and missing years in the accepted subset in Figure 1) to justify the conclusion 'TLA+ usage has surged since 2015.' The statement that the trend 'is likely to continue' is a prediction unsupported by either the regression or the qualitative data. At minimum, the RQ1 answer should be restricted to 'publications matching the search query increased,' with the industrial-adoption claim presented as a qualitative observation from the 16 case studies.","section":"Section 3.1 and Figure 1"}],"minor_comments":[{"comment":"The stacked bar chart is difficult to read; please add a table with the exact per-year counts for each category (optimistic, deduplicated, rejected, low-affinity, high-affinity).","section":"Figure 1"},{"comment":"The sentence 'which potentially skews our trend analysis and leading to misleading conclusions' should be rephrased, e.g., 'which could skew our trend analysis and lead to misleading conclusions.'","section":"Section 4"},{"comment":"The phrase 'This insight of a growing overall trend' should be 'This observation of a growing overall trend.'","section":"Section 3.1"},{"comment":"The exclusion criterion 'no access to paper over university network' may introduce accessibility bias; please discuss this in the threats to validity section.","section":"Section 2.2"},{"comment":"No inter-rater reliability measure (e.g., Cohen's kappa) is reported for the screening and affinity-scoring phases; please add such a measure or justify its absence.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope for an SE/formal-methods venue and the Zenodo dataset is a useful artifact. However, the RQ1 trend claim and the conclusion overstate the evidence; the authors must either add a proper baseline and analyze the 16-paper corpus, or substantially weaken the claims. The remaining qualitative findings are salvageable and likely publishable after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Roman and colleagues have put together the first systematic review I know of dedicated to industrial TLA+ practice. The protocol is standard Kitchenham/Petersen, the dataset is on Zenodo, and the 16 high-affinity papers are a genuinely useful curated list. The qualitative answers to RQ2–RQ4 are solid: cloud dominance, early-design/debugging emphasis, reported benefits in bug-finding and design understanding, and the two recurring challenges (learning curve, abstraction choice). This will be a handy entry point for practitioners considering TLA+ and for researchers looking for reported pain points.\n\nThe soft spot is RQ1. The 'significant upward trend' is computed on the 261 deduplicated Google Scholar hits, which the authors themselves concede in Section 4 'may include irrelevant or nonsensical papers.' That set is a measure of publication volume matching a broad search string, not of industrial usage. The 16 high-affinity papers are too few and too clumpy for a per-year trend. So the conclusion that 'TLA+ usage has surged since 2015' is not directly supported by the data presented. The authors do flag this in their threats to validity, which is to their credit, but the abstract and conclusion still state the surge as a finding. That needs to be reworded to something like 'publications matching our search criteria have increased,' or the trend analysis needs a baseline comparison (e.g., TLA+ papers without industry terms) to make the industry-specific claim.\n\nThe success-story bias in RQ3/RQ4 is acknowledged but not deeply addressed; a sensitivity analysis or a discussion of how the 16 papers were selected for positive reporting would strengthen it. That is a minor issue, not a fatal one.\n\nOverall, this is a careful, reproducible review with an honest limitations section. The central qualitative findings are credible. I would accept it with conditions: fix the RQ1 framing and soften the conclusion. The dataset and curated bibliography alone are worth having.","headline":"A useful, transparent SLR on industrial TLA+ that overreaches in its RQ1 trend claim; the qualitative synthesis is the real contribution.","tokens_in":8847,"tokens_out":2132,"would_cite":true,"duration_ms":19990,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of a decade of industry reports finds TLA+ adoption surging since 2015, led by cloud systems.","keywords":["TLA+","PlusCal","formal methods","systematic literature review","industrial adoption","model checking","cloud computing","software engineering"],"falsifier":"If a re-run of the same Google Scholar search string over the same period produced a flat or declining deduplicated count, or if the linear regression on the 16 high-affinity papers alone (rather than the 261 deduplicated papers) had $p \\ge 0.05$, the paper's 'significant upward trend' and 'surge since 2015' claims would not be supported. The paper itself reports the regression only for the 261-paper set and notes that set may include irrelevant or nonsensical papers.","tokens_in":7828,"feed_emoji":"📈","tokens_out":7159,"duration_ms":63064,"temperature":0.7,"pith_summary":"The paper sets out to establish, through a systematic review of the past decade, that TLA+—a formal specification language for describing and model-checking system behavior—has moved into real industrial use, and that this use is growing and concentrated in specific settings. It argues that industry reports since 2015 show adoption surging, particularly for cloud systems, where TLA+ is used mostly during early design and debugging rather than during implementation. The review claims these reports confirm TLA+'s promised benefits: finding subtle bugs such as deadlocks and race conditions and improving system understanding, while also surfacing two persistent obstacles, a steep learning curve and difficulty choosing the right abstraction level. A sympathetic reader would care because the findings give practitioners a curated evidence base for deciding whether to adopt TLA+ and give researchers a map of where tooling and training should improve.","feed_headline":"Industry TLA+ use has surged since 2015, review finds","feed_subtitle":"A decade of industry reports shows cloud systems leading adoption, with early-design bug finding the main payoff.","key_machinery":"The review mechanism is a three-stage systematic literature review: an optimistic Google Scholar search with a keyword string combining TLA+ and PlusCal with practice and industry terms, followed by a reduction pipeline of deduplication (290 to 261 papers), rejection (261 to 37), and an affinity filter scoring each paper on reasons for using formal methods, modeling assumptions, the model-implementation link, and practical drawbacks (37 to 16 high-affinity papers). A linear regression on the 261 deduplicated papers supplies the trend answer for RQ1, while the 16 high-affinity papers provide the qualitative evidence for application settings, benefits, and challenges.","core_discovery":"The paper claims that TLA+, a temporal-logic specification language with the PlusCal front-end, has moved from a niche academic tool to a growing industrial practice over 2013-2023. Analyzing 290 initially found publications and, after filtering, 16 high-affinity industry reports, it concludes that TLA+ usage has surged since 2015, is most common in the cloud industry (63% of the included papers), and is applied mainly during early design and debugging. The reports consistently credit TLA+ with finding subtle bugs and improving system understanding, while naming a steep learning curve and abstraction-level choice as the main barriers. The review frames the Amazon Web Services report [10] as a catalyst because 44% of the included papers cite it as motivation. The overall upward trend is quantified by a linear regression on the 261 deduplicated papers with slope $S = 3.67$ and $p < 0.001$.","pith_inferences":["Beyond the paper: if the observed cross-company contagion holds, formal-specification skills could become a visible differentiator in cloud-infrastructure hiring within a few years.","Beyond the paper: the 63% cloud concentration suggests that the next high-impact TLA+ targets are cloud control planes and consensus protocols, where the state space is bounded enough for exhaustive model checking.","Beyond the paper: the Google-Scholar-only corpus cannot see proprietary or defunct industrial work, so the true adoption curve may be steeper or shallower than reported; interviews, surveys, or a second search engine would test which.","Beyond the paper: applying the same affinity filter to other formal methods such as Alloy, B, or VDM would show whether the post-2015 surge is TLA+-specific or part of a broader formal-methods revival."],"forward_implications":["Industrial TLA+ use is concentrated in cloud computing, with more than 60% of the high-affinity papers in that domain.","TLA+ is mostly used during early design (81% of label occurrences), then debugging (44%), and least during implementation (38%).","Practitioners report that TLA+ finds subtle bugs such as deadlocks, race conditions, and stack overflow errors, and that it improves overall system understanding.","Adoption is hindered by a steep learning curve and by difficulty choosing the right abstraction level; PlusCal is seen as easing the learning curve.","The significant upward trend in TLA+ publications is likely to continue because success stories such as the Amazon Web Services report are cited as motivators by later adopters."],"supporting_citations":[{"why":"Supplies the flagship cloud-provider success story; the review identifies it as the catalyst because 7 of the 16 included papers cite it as motivation.","marker":"[10]"},{"why":"Provides the systematic-literature-review guidelines that structure the entire review process.","marker":"[21]"},{"why":"Provides mapping-study guidelines used for collecting, filtering, and analyzing the literature.","marker":"[24]"},{"why":"Supplies the how-to guide behind the review protocol, including inclusion and exclusion criteria, quality assessment, and data extraction.","marker":"[18]"},{"why":"Gives the justification for relying on Google Scholar and grey literature to capture industrial reports, the basis of the search corpus.","marker":"[28]"},{"why":"Exemplifies the bug-finding benefit and the broader trend of companies adopting the practice.","marker":"[2]"},{"why":"Exemplifies the cloud database application domain and explicitly cites the Amazon report [10] as the key motivator.","marker":"[3]"},{"why":"Used to caution that adoption may follow hype cycles rather than a steady linear increase.","marker":"[20]"},{"why":"Used alongside the hype-cycle literature to caution that the upward trend may be nonlinear.","marker":"[27]"}],"fun_headline_variants":["TLA+ industrial use surged since 2015, review finds","Cloud leads industrial TLA+ surge: 63% of reports","TLA+ catches bugs early, but steep learning curve persists","Industry TLA+ reports triple in decade, led by cloud"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's trend and surge conclusions rest on the assumption that the 261 deduplicated Google Scholar matches—which the authors concede may include irrelevant or nonsensical papers—accurately track industrial TLA+ activity rather than academic output or search-engine noise.","fun_headline_variants_meta":{"raw":{"variants":["TLA+ industrial use surged since 2015, review finds","Cloud leads industrial TLA+ surge: 63% of reports","TLA+ catches bugs early, but steep learning curve persists","Industry TLA+ reports triple in decade, led by cloud"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2548,"prompt_tokens":826,"completion_tokens":1722,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":1649}},"tokens_in":442,"tokens_out":1722,"duration_ms":11939,"temperature":1.0,"reasoning_tokens":1649,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:56:46.483061+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a re-run of the same Google Scholar search string over the same period produced a flat or declining deduplicated count, or if the linear regression on the 16 high-affinity papers alone (rather than the 261 deduplicated papers) had $p \\ge 0.05$, the paper's 'significant upward trend' and 'surge since 2015' claims would not be supported. The paper itself reports the regression only for the 261-paper set and notes that set may include irrelevant or nonsensical papers.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the systematic-literature-review guidelines that structure the entire review process."},{"cited_title":"In: Huisman, M., Păsăreanu, C., Zhan, N","cited_arxiv_id":null,"evidence_quote":"Exemplifies the cloud database application domain and explicitly cites the Amazon report [10] as the key motivator."},{"cited_title":"Technological Forecasting and Social Change108, 28–41 (Jul 2016)","cited_arxiv_id":null,"evidence_quote":"Used to caution that adoption may follow hype cycles rather than a steady linear increase."},{"cited_title":"Social science, Free Press, New York Lon- don Toronto Sydney, 5 edn","cited_arxiv_id":null,"evidence_quote":"Used alongside the hype-cycle literature to caution that the upward trend may be nonlinear."}],"review_version":1}