{"id":"001cba40-4c50-42d8-b304-1ac64391e6d8","arxiv_id":"1906.11940","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Empirical study of Stack Overflow posts on ML libraries finds prevalent API misuses, lack of early error detection, and a need for more SE research on debugging and model behavior understanding.","lead":"This paper manually classifies 3,243 Stack Overflow questions about ten ML libraries into seven pipeline stages and identifies common difficulties. The results point to missing static analysis tools and API design issues that could inform better SE support for ML.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of 3,243 highly-rated SO posts to general ML library difficulties is unverified and central to the extrapolation","rationale":"The reader's weakest assumption matches the load-bearing step exactly; the study design itself (SO-only, highly-rated filter, manual classification) contains no independent check on external validity, so the concern is already correctly located and the low-confidence UNVERDICTED verdict does not require adjustment.","tokens_in":1745,"tokens_out":276,"duration_ms":12144,"concrete_test":"Sample 500 recent GitHub issues from the ten libraries' repositories, classify them into the same seven pipeline stages using the paper's scheme, and test whether the stage distribution differs from the SO set by more than 15 percentage points in any stage (chi-square or proportion test).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (urgent need for static/dynamic analyses and API design improvements) is derived from classifying the selected posts into pipeline stages and observing patterns such as prevalent API misuses. This requires that the highly-rated SO subset is representative of developer difficulties; the paper performs no cross-validation against GitHub issues, developer surveys, or usage telemetry. Highly-rated posts may over-represent questions that attract answers while under-representing silent failures or internal-tooling problems.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports a manual examination and statistical analysis of 3,243 highly-rated Stack Overflow Q&A posts across ten ML libraries (TensorFlow, Keras, scikit-learn, etc.). Posts are classified into seven stages of a typical ML pipeline; the authors then address four research objectives on the most difficult stage, nature of problems, library differences, and temporal stability, concluding that static/dynamic analyses are absent, API misuses are prevalent, and API design improvements plus further SE research are urgently needed.","tokens_in":1844,"tokens_out":497,"duration_ms":17298,"significance":"If the classification process proves reliable and the highly-rated SO sample is representative, the work supplies concrete evidence of tooling gaps at the SE-ML boundary and could usefully guide priorities for static analysis, debugging support, and API usability research. The multi-library scope and pipeline-stage framing are strengths that would make the findings actionable for both researchers and library maintainers.","major_comments":[{"comment":"Methodology section (data collection and classification): the abstract and text describe manual examination and assignment to seven pipeline stages but supply no information on inter-rater reliability, how the seven stages themselves were validated or pilot-tested, or the precise exclusion criteria applied to arrive at the final 3,243 posts. These omissions directly affect the soundness of every subsequent statistical claim and the identification of 'most difficult' stages.","section":"Methodology (data collection and classification)"},{"comment":"Results and Discussion sections: the central extrapolation that 'both static and dynamic analyses are mostly absent and badly needed' and that 'API misuses are prevalent' rests on the assumption that the selected highly-rated SO posts represent the difficulties faced by developers in general. No cross-validation against GitHub issues, developer surveys, or usage telemetry is reported, leaving the generalization load-bearing for the 'urgent need' conclusion.","section":"Results and Discussion"}],"minor_comments":[{"comment":"Abstract: the phrase 'highly-rated' is used without stating the exact rating threshold or vote count applied during selection.","section":"Abstract"},{"comment":"The description of the four research objectives would benefit from explicit mapping to the statistical tests or tables that address each one.","section":"Introduction / Research Objectives"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the two major comments below, indicating planned revisions where appropriate.","responses":[{"response":"We agree that the methodology section would benefit from greater transparency. The seven pipeline stages were derived from standard descriptions in the ML and SE literature (e.g., data preparation, model training, evaluation). Exclusion criteria included posts tagged with the ten libraries, having an accepted answer, and a minimum score threshold to focus on highly-rated content; non-English posts and duplicates were removed. Classification was performed by the first two authors, with disagreements resolved via discussion until consensus. No formal inter-rater reliability statistic (e.g., Cohen's kappa) was computed. We will expand the methodology section with explicit stage definitions, a description of the pilot phase used to refine the stages, the exact exclusion rules, and the consensus process.","revision_made":"yes","referee_comment":"Methodology section (data collection and classification): the abstract and text describe manual examination and assignment to seven pipeline stages but supply no information on inter-rater reliability, how the seven stages themselves were validated or pilot-tested, or the precise exclusion criteria applied to arrive at the final 3,243 posts. These omissions directly affect the soundness of every subsequent statistical claim and the identification of 'most difficult' stages."},{"response":"The study is explicitly scoped to highly-rated Stack Overflow posts, which serve as a public record of developer difficulties that have been vetted by the community through votes and answers. This source is commonly used in empirical SE research on API usage and learning barriers. We acknowledge that the absence of triangulation with GitHub issues or surveys limits the strength of broad generalizations. We will revise the discussion and threats-to-validity sections to (a) more precisely bound the claims to the SO dataset and (b) explicitly call for future multi-source validation studies.","revision_made":"partial","referee_comment":"Results and Discussion sections: the central extrapolation that 'both static and dynamic analyses are mostly absent and badly needed' and that 'API misuses are prevalent' rests on the assumption that the selected highly-rated SO posts represent the difficulties faced by developers in general. No cross-validation against GitHub issues, developer surveys, or usage telemetry is reported, leaving the generalization load-bearing for the 'urgent need' conclusion."}],"tokens_in":1437,"tokens_out":504,"duration_ms":26800,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This study manually classifies over three thousand Stack Overflow posts across Tensorflow, Keras, scikit-learn and seven other libraries into seven ML pipeline stages. It reports distributions, some time trends, and notes that API misuse questions are common while static and dynamic analysis support is thin. That classification is the actual new data point; earlier SO studies were broader and did not drill into these specific libraries and stages with the same granularity.","headline":"The paper catalogs questions on ten ML libraries from highly-rated Stack Overflow posts and breaks them down by pipeline stage, but the push for urgent new tooling rests on untested assumptions about what those posts represent.","tokens_in":2302,"tokens_out":169,"would_cite":false,"duration_ms":16800,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical SE study of ML library Q&A on SO has zero overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper performs manual classification of 3243 SO posts into ML pipeline stages (data prep, modelling, training, etc.) and reports prevalence of API misuses and need for static/dynamic analyses. This is standard empirical software engineering; it invokes no recognition cost J(x), golden-ratio identities, 8-tick periodicity, Alexander duality for D=3, or any theorem from the RS Lean corpus (AbsoluteFloorClosure, AlexanderDuality, Cost.FunctionalEquation, etc.). Domain is orthogonal to the RS foundation-to-constants chain.","tokens_in":57871,"confidence":"high","tokens_out":143,"duration_ms":5414,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Analysis of 3,243 Stack Overflow posts on ten ML libraries shows static and dynamic analyses are absent and API misuses are common.","keywords":["machine learning libraries","stack overflow","developer questions","API misuse","ML pipeline","software engineering","error detection"],"falsifier":"A follow-up survey or interview study of practicing ML developers that finds their most common problems do not match the distribution of stages and error types identified in the Stack Overflow posts.","tokens_in":2661,"feed_emoji":"","tokens_out":608,"duration_ms":23397,"temperature":0.7,"pith_summary":"The paper examines 3,243 highly-rated questions from Stack Overflow about ten ML libraries to map the difficulties developers encounter when incorporating machine learning into systems. Questions are classified into seven stages of a standard ML pipeline, then analyzed statistically across four objectives that cover the hardest stages, problem types, library differences, and changes over time. The results indicate that support for early error detection is lacking and that API design changes would address frequent misuses. The work concludes that software engineering research must address these gaps to help developers avoid problems during model training and evaluation.","feed_headline":"Stack Overflow study finds ML libraries lack error detection tools","feed_subtitle":"3,243 posts across ten libraries show absent static analyses and prevalent API misuses during training and evaluation.","key_machinery":"Manual classification of questions into seven stages of an ML pipeline followed by statistical analysis across libraries and time periods.","core_discovery":"Our findings reveal the urgent need for software engineering research in this area. Both static and dynamic analyses are mostly absent and badly needed to help developers find errors earlier. API misuses are prevalent and API design improvements are sorely needed. Last and somewhat surprisingly, a tug of war between providing higher levels of abstractions and the need to understand the behavior of the trained model is prevalent.","pith_inferences":["Library maintainers could instrument their code with additional checks at the training and evaluation stages where questions cluster.","The observed tension between abstraction and model transparency may affect adoption rates of newer high-level ML frameworks.","Educational materials and documentation for ML libraries should prioritize the pipeline stages that generate the most questions."],"forward_implications":["Static and dynamic analysis techniques must be developed specifically for ML library usage to catch errors before runtime.","Debugging support for ML systems requires substantially more research attention.","Redesign of ML library APIs is needed to reduce the rate of misuses observed in the questions.","Approaches that reconcile high-level abstractions with visibility into trained model internals should be explored."],"fun_headline_variants":["Large SO study shows ML libs miss error detection","API misuses dominate ML library Stack Overflow posts","ML model understanding conflicts with abstractions","Study reveals need for ML library SE research"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 3,243 highly-rated Q&A posts selected from Stack Overflow are representative of the difficulties faced by software developers when learning about and using ML libraries in their systems.","fun_headline_variants_meta":{"raw":{"variants":["Large SO study shows ML libs miss error detection","API misuses dominate ML library Stack Overflow posts","ML model understanding conflicts with abstractions","Study reveals need for ML library SE research"]},"model":"grok-4.3","cost_usd":0.006167,"raw_usage":{"total_tokens":2841,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":61665500,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2093,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":53,"duration_ms":26928,"temperature":1.0,"reasoning_tokens":2093,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T14:13:31.035028+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A follow-up survey or interview study of practicing ML developers that finds their most common problems do not match the distribution of stages and error types identified in the Stack Overflow posts.","supporting_citations":[],"review_version":1}