{"id":"347c78bd-a70b-4f1c-a63b-2647071cec1a","arxiv_id":"1908.01502","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In one Finnish micro-company, the most significant technical debt was a supplier disagreement and missing test automation, driven mainly by budget and time pressure.","lead":"A two-hour focus group at a small Finnish software company found that its biggest technical debts came from a dispute with an outside supplier and from missing automated tests. The study is a compact real-world example of how budget and deadline pressure make small teams defer work that later costs more to fix.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table I's classification and vote totals are internally inconsistent; the claimed rankings of Test TD and lack of test automation cannot be recovered from the reported data.","rationale":"The reader's verdict was CONDITIONAL, identifying retrospective self-report bias as the weakest assumption. I agree that self-report bias is a real limitation, but the more immediate and more decisive problem is internal: the paper's own data tables cannot support the central ranking because of arithmetic and classification inconsistencies. The article's contribution is precisely the claim that Test TD and Requirements TD are the most significant types, and that supplier disagreement and lack of test automation are the most significant items. These claims are tied to vote totals in Table I. If TD5 is classified as Infrastructure, as the text states, the Test TD total drops from 11 to 4 and Infrastructure rises to 9, reversing the type-level ranking. The grand-total mismatch (30 in Table 1a, 33 in Table 1b) and the unexplained discrepancy with the 50 available votes further indicate that the reported numbers are not reliable. This is not a matter of disagreeing with the qualitative conclusion; it is a matter of the evidence presented being internally inconsistent. A corrected table could potentially restore the conclusion, but as published the central empirical result is not verifiable from the paper. Hence REJECT rather than CONDITIONAL or UNCHANGED. I credit the authors for listing threats to validity and for framing the study as a single case, but those strengths do not repair the arithmetic and classification contradictions in the principal results table.","tokens_in":8486,"tokens_out":3584,"duration_ms":37223,"concrete_test":"Reconstruct the full contingency table from the original focus-group artifacts (post-it notes and dot votes) or ask the authors to provide it: assign each TD item to exactly one TD type using Li et al.'s definitions, list the vote count per item, and verify that item sums, type sums, and grand total all agree. In particular, decide whether TD5 is Test TD or Infrastructure TD. If TD5 is Infrastructure, the abstract's 'Test TD most significant' claim is false and the top-ranked type is Requirements TD alone; if TD5 is Test TD, the item's label and the type table must be corrected and the total votes must be reconciled with the stated five participants times ten votes (50) versus the printed 30.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central empirical result, that Test and Requirements TD are the most significant types and that supplier disagreement and lack of test automation are the top debt items, rests on Table I in Section IV. That table contains contradictions that undermine the headline ranking. First, the procedure says each of five participants received ten votes, so up to 50 votes were available, yet Table 1a sums to 30 and no explanation is given for missing votes. Second, the type-level table (Table 1b) sums to 33 while its stated Grand Total is 30. Third, and most damaging, TD5 'Lack of automatic testing' is explicitly labeled 'Infrastructure TD' in the text, but the Test TD total of 11 is only reachable if TD5's 7 votes are counted under Test TD. If TD5 is classified as Infrastructure, as its label states, then Test TD totals 4 (TD6 + TD7) and Infrastructure TD totals 9 (TD3 + TD5), so the abstract's assertion that 'Specification and test TD are the most significant types of TD' is no longer supported: Requirements TD (10) would be the only top type, and the basis for ranking TD5 as co-most-significant disappears. The cause-count table also records budget and time as tied at 5, but the narrative lists multiple causes for individual items, so the root-cause ranking is likewise not reproducibly derived from the reported data. Because the abstract's strongest claims depend directly on these sums and classifications, the internal inconsistency is a load-bearing flaw, not a stylistic one.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory case study of technical debt (TD) in a Finnish micro-enterprise developing a SaaS sales-channel management tool. The authors conducted a two-hour focus group with five company members (CTO, CFO, CMO, and two developers), moderated by a co-author. Participants listed postponed activities, classified them into TD types using a 10+1 category scheme, identified causes and effects, and voted on the most harmful items. The authors conclude that supplier disagreement and lack of test automation are the most significant TD items, that Test and Requirements TD are the most significant types, and that budget and time constraints are the main root causes. The paper also proposes mitigation strategies: learning from customers, careful estimation, and continuous improvement.","tokens_in":8724,"tokens_out":8074,"duration_ms":68423,"significance":"If the data are reliable, the study offers a rare qualitative account of TD in a small company, showing that non-code-level debt (requirements conflicts, missing test automation, and supplier issues) can dominate code-level shortcuts. The focus-group design with mixed roles is appropriate for surfacing organizational causes, and the authors explicitly acknowledge several threats to validity, including management presence and the two-hour time limit. The paper reports raw vote counts and quotes participant statements, which adds transparency. However, the significance is reduced by the small number of participants, the lack of inter-rater reliability checks or independent coding, and the absence of outcome data linking postponed activities to actual costs. As a single-case exploratory study, the contribution is modest but potentially reusable for future comparisons if the data presentation is corrected.","major_comments":[{"comment":"The type-level point totals are inconsistent with the item-level data. The sum of the values in Table I(b) is 33, while the table's stated Grand Total is 30; this appears to result from counting TD1 (3 votes) under both Requirements TD and Organizational TD, but the table does not say so. More seriously, TD5 is defined in Section IV.A as 'Lack of automatic testing costs more in the future (Infrastructure TD)', yet the Test TD total of 11 in Table I(b) is only attainable if TD5's 7 votes are counted under Test TD. If TD5 is classified as Infrastructure TD as labeled, then Test TD totals 4 (TD6+TD7) and Infrastructure TD totals 9 (TD3+TD5), which would contradict the abstract's claim that 'Specification and test TD are the most significant types of TD'. The authors must either correct the TD5 classification, revise the type totals, or qualify the conclusion accordingly.","section":"Section IV, Table I(b) and Section IV.A"},{"comment":"The cause-count table reports 'Design issues' with a count of 1, but no TD item in the text identifies design issues as a cause; for instance, TD4 (Design TD) is attributed only to Time constraints. The mapping between the TD items and the cause counts in Table I(c) is not reproducible from the text. The authors should provide the explicit mapping from each TD item to its causes or reconcile Table I(c) with the item descriptions.","section":"Section IV, RQ2 and Table I(c)"},{"comment":"The voting procedure states that each of the five participants received ten votes, which implies up to 50 votes, but Table I(a) sums to 30. The paper does not explain whether participants were allowed to use fewer than ten votes or whether some votes were invalid. If all ten votes per participant were expected, the missing 20 votes could bias the ranking of the most significant TD items; if abstentions were permitted, this should be stated explicitly and considered in the interpretation.","section":"Section III.B (T6) and Table I(a)"}],"minor_comments":[{"comment":"'As reported in Table 1a' should reference Table 1b, since the type-level totals appear in Table 1b.","section":"Section IV, RQ1"},{"comment":"The sentence beginning 'The testing budget was too low ...' contains a fragment ('since the company did not even have enough time to the prioritization of the features and tasks...') that appears to be a copy-paste error from TD1.","section":"Section IV.A, TD5"},{"comment":"Table I(a) lists 'TD0. Technical shortcuts', but the text defines 'TD9 Technical shortcuts (Code TD)'; the numbering is inconsistent.","section":"Section IV.A, Table I(a)"},{"comment":"The company is described as 'a micro-enterprise (less than 10 person)', while the title and abstract use 'SME'; the terminology should be harmonized.","section":"Section III"},{"comment":"The caption states that 'Motivations are counted once for each TD', but Table I(b) appears to count TD1 twice (under Requirements and Organizational TD); the caption or the table should be corrected.","section":"Table I caption"},{"comment":"The Data Analysis subsection gives no detail on how the TD items were assigned to the eleven categories or whether multiple participants independently validated the classifications; a sentence on this would help the reader assess the reliability of the type-level totals.","section":"Section III.C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a single-case exploratory study; its value depends on the trustworthiness of the reported data. The internal inconsistencies in Table I are the main obstacle to acceptance; with a corrected and internally consistent data presentation, the paper could make a useful, if modest, contribution to the TD case-study literature. I would also flag that several self-citations in the related work and future work sections are only tangentially related to the study's evidence chain; the authors might consider pruning them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nRead the Finnish SME technical debt case study. Short version: it's a plausible, well-intentioned qualitative case, but the headline numbers don't survive a re-tally. The abstract says test and requirements debt dominate, with supplier disagreement and lack of test automation as the top items. If you take the tables as printed, that ranking is not recoverable.\n\nWhat's actually new: a micro-enterprise case where the biggest perceived debts are a supplier conflict over MVP scope and missing automated testing, with concrete quotes from the CTO and CFO. The mitigation suggestions (prototyping, careful estimation, continuous improvement) are practical, and the authors list the main validity threats, including the presence of senior management in the focus group and the two-hour limit. That is honest.\n\nThe problem is Table I. Five participants had ten votes each, so up to 50 votes were available, but the item table sums to 30 and the type table sums to 33. The reason is a double count: TD1 is counted under both Requirements and Organizational TD, adding the same 3 points twice. And TD5, explicitly labeled Infrastructure TD in the text, is counted as Test TD in the type table. If you move TD5 back to Infrastructure, Test TD drops from 11 to 4 and Infrastructure rises from 2 to 9, so the abstract's claim that 'specification and test TD are the most significant' no longer holds. The cause table also has a 'Design issues' entry with no corresponding cause listed in the narrative. These aren't cosmetic errors; they're the evidence base for the paper's central claims.\n\nThe related-work and citation pattern is fine. There are many self-citations, but mostly in the context of prior TD taxonomies and future tool suggestions; the case narrative doesn't depend on them.\n\nWho this is for: practitioners in small software companies, and researchers looking for qualitative TD data points. The paper is a legitimate case study worth a serious referee, but it needs major revision. The authors should publish raw vote tallies, clarify how each item was classified, fix the table arithmetic, and make the abstract match the corrected numbers. I'd recommend send for review, but with a strong note that the current tables are internally inconsistent.\n\nLet me know if you want to discuss.","headline":"A useful qualitative case study whose headline numbers don't add up: Table I's vote sums and type classifications are internally inconsistent, and the abstract's ranking of Test and Requirements debt doesn't survive a re-tally.","tokens_in":9291,"tokens_out":4790,"would_cite":false,"duration_ms":44964,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Finnish micro-enterprise's worst technical debt was a supplier dispute and missing automated tests.","keywords":["technical debt","small and medium enterprise","focus group","requirements debt","test debt","root causes","minimum viable product","case study"],"falsifier":"Go to the company's records and check the two highest-ranked items directly: inspect the version history and test suite to see whether automated tests were really absent during the period the participants described, and interview the external supplier to see whether the MVP disagreement and its costs are corroborated. If the supplier disputes the account, or if test automation existed earlier than claimed, the severity ranking loses its factual basis. A second check would be a replication in the same company using anonymous individual interviews rather than a group session; a different top-two ranking would show that the focus-group setting shaped the result.","tokens_in":8276,"feed_emoji":"💸","tokens_out":6862,"duration_ms":67342,"temperature":0.7,"pith_summary":"This paper reports a two-hour focus group held inside a Finnish micro-enterprise that builds a SaaS sales-channel product, with the CTO, CFO, CMO, and two developers recalling which activities they had postponed and what happened as a result. The authors are trying to establish that this company's most significant technical debt was not code-level shortcuts but two non-code items: a disagreement with an external supplier over the Minimum Viable Product, and the absence of automated testing. They also claim that specification and test debt were the dominant types of debt, and that budget and time constraints were the most frequent root causes. A sympathetic reader would care because the study offers a concrete, role-diverse account of how debt arises in a small company and suggests that learning from customers, careful estimation, and continuous improvement are the levers for avoiding it. The paper is explicit that not every postponement became debt: some delays met deadlines cheaply, while others accrued interest out of proportion to the benefit.","feed_headline":"Supplier clash and absent test automation top SME debt","feed_subtitle":"A five-person focus group ranked a supplier MVP fight and missing automated tests above code shortcuts.","key_machinery":"The load-bearing mechanism is the structured focus group combined with a classification scheme and a voting procedure. Participants first wrote postponed activities on post-it notes, then assigned each activity to one of eleven technical-debt categories (ten from the cited taxonomy plus a new Organizational Debt category), then attached causes using the 5-Whys technique. Severity was measured by dot-voting: each of the five participants had ten adhesive dots to place on the problems they considered most harmful, producing the 30-point total that drives the ranking. The moderation was deliberately hands-off, with an author who did not vote or propose items, and the whole session was limited to about two hours.","core_discovery":"On the paper's own terms, the central discovery is an empirical ranking produced by the company's own staff: the most harmful technical debt items were a disagreement with the supplier about what the Minimum Viable Product should contain, and the lack of automatic testing, each receiving 7 of the 30 voting dots. Grouped by type, test debt received 11 points and requirements debt 10, well above code debt, which received 5. The root causes named most often were budget constraints and time constraints, each attached to five of the ten debt items, followed by estimation issues with three. The authors interpret this as evidence that postponement is not automatically debt: it becomes debt when the deferred work later costs more to repair than it bought in deadline relief, and the company's costly episodes were tied to weak validation of requirements, underestimated testing effort, and supplier relations.","pith_inferences":["The authors do not say this, but their data suggest a coupling: missing automated tests made requirement changes riskier, so the supplier's refusal to iterate became more expensive than it would have been with a safety net.","A testable extension of the root-cause ranking is that companies using prototype-based validation with customers before contracting will report lower requirements debt; this could be checked with a multi-company focus group replication.","The two-hour, mixed-hierarchy focus group design may systematically undercount debt that developers are reluctant to admit in front of the CTO; an anonymous survey or diary study could reveal a different distribution.","If the pattern generalizes beyond this firm, technical debt dashboards for SMEs should include requirement and test debt indicators, not only code smells and coverage metrics, which aligns with the paper's own call for continuous quality monitoring."],"forward_implications":["If these results hold for the case company, small software firms should expect requirement and test debt, rather than code debt, to be their most expensive backlog items.","The dominance of budget and time constraints as causes implies that better estimation and earlier customer validation are the most direct mitigation levers, matching the paper's three recommendations.","The finding that not all postponement creates debt suggests that companies should judge deferral by whether it accrues future interest, not by the mere existence of shortcuts.","The supplier-disagreement item implies that outsourcing the first version of a product shifts technical debt risk to the requirements boundary, where contractually fixed specifications can block iterative learning.","Because test debt received the highest type-level score, investing in automated testing early is presented as a way to reduce both regression risk and the manual effort that later consumes developer time."],"supporting_citations":[{"why":"Defines technical debt as postponed software maintenance activities taken for short-term payoff, the concept the whole study measures.","marker":"[2]"},{"why":"Supplies the ten-category technical debt taxonomy the focus group used to classify each postponed activity by debt type.","marker":"[4]"},{"why":"Provides the 5-Whys technique used to analyze the root causes behind each postponement.","marker":"[8]"},{"why":"Supplies the Minimum Viable Product concept that defines the supplier disagreement ranked as the most significant debt item.","marker":"[9]"}],"fun_headline_variants":["Supplier row and missing tests rank as SME's top debt","Test and requirements debt beat code in Finnish SME study","SME staff: time and budget constraints drive worst debt","SME's costliest tech debt from supplier fight, no test automation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking depends on five employees accurately remembering, in one two-hour session with the CTO, CFO, and CMO present, which postponed activities later became costly and why; if those memories or answers were biased, the top-ranked debt items would not be the true ones.","fun_headline_variants_meta":{"raw":{"variants":["Supplier row and missing tests rank as SME's top debt","Test and requirements debt beat code in Finnish SME study","SME staff: time and budget constraints drive worst debt","SME's costliest tech debt from supplier fight, no test automation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1304,"prompt_tokens":898,"completion_tokens":406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":337}},"tokens_in":514,"tokens_out":406,"duration_ms":5099,"temperature":1.0,"reasoning_tokens":337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:10:45.556335+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Go to the company's records and check the two highest-ranked items directly: inspect the version history and test suite to see whether automated tests were really absent during the period the participants described, and interview the external supplier to see whether the MVP disagreement and its costs are corroborated. If the supplier disputes the account, or if test automation existed earlier than claimed, the severity ranking loses its factual basis. A second check would be a replication in the same company using anonymous individual interviews rather than a group session; a different top-two ranking would show that the focus-group setting shaped the result.","supporting_citations":[{"cited_title":"Investigating architectural technical debt accumulation and refactoring over time: A multiple-case study,","cited_arxiv_id":null,"evidence_quote":"Defines technical debt as postponed software maintenance activities taken for short-term payoff, the concept the whole study measures."}],"review_version":1}