{"id":"cfe38579-4b8b-4d73-885a-d66ca72bd15b","arxiv_id":"2607.06043","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Survey data from 1,134 MSI students shows positive STEM recognition experiences predict STEM identity, but do not fully predict STEM career aspirations.","lead":"This study surveyed 1,134 students at Minority Serving Institutions to test how recognition experiences (e.g., teacher recommendations, STEM awards) correlate with STEM identity and career aspirations. It finds positive recognition predicts STEM identity but only partially predicts career aspirations, suggesting interventions need to be more targeted.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The 31% listwise deletion is the most load-bearing concern; if missingness is related to recognition experiences or outcomes, the reported associations could shift substantially under multiple imputation.","rationale":"The reader identified recall bias as the weakest assumption, which is a valid concern for any retrospective survey study. However, I find the 31% listwise deletion to be the more quantitatively load-bearing issue because (1) it directly shapes the analytic sample composition, (2) the direction of bias is predictable and potentially inflates the key associations, (3) it is immediately testable via the multiple imputation the authors already plan, and (4) the null finding on misrecognition—which the reader noted as interesting—could be an artifact of the missingness pattern rather than a genuine result. The recall bias concern is real but more fundamental and less immediately resolvable; it would require a longitudinal design to fully address. The missing data concern is the one that could change the verdict on the current data. That said, the reader's CONDITIONAL verdict is appropriate: the findings are worth disseminating as preliminary but require the planned imputation and fuller reporting before they can be considered robust. My concern does not move the verdict because the reader already accounted for the missing data issue in their rationale, even though they designated recall bias as the weakest assumption.","tokens_in":7350,"tokens_out":3452,"duration_ms":255766,"concrete_test":"Implement the planned multiple imputation (e.g., mice with 20+ imputations) using all auxiliary variables available in the NSF survey. Re-estimate both the logistic regression (career aspirations) and linear regression (STEM identity). If the recognition β coefficients shift by more than ~20% or any p-value crosses 0.05—particularly for the misrecognition items, which are currently null—the current findings are not robust to the missingness assumption and the headline claims would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The study removes 354 of 1,134 observations (31%) via listwise deletion. The authors acknowledge this and plan multiple imputation, but the current results rest on the assumption that missingness is completely at random (MCAR). This is implausible: students who had fewer recognition experiences may have been more likely to skip those survey items (having nothing to report reduces item engagement), which would make the analytic sample over-represent recognized students and inflate the recognition–identity and recognition–aspiration associations. The null result for misrecognition is equally vulnerable—if students who experienced misrecognition were differentially likely to leave those items blank, the null may reflect missingness rather than a true absence of association. With 31% missingness, even moderate departures from MCAR could materially change the reported β coefficients and p-values. The PCA also raises a secondary concern: RC1 and RC2 explain only 43% of total variance (27% + 16%), suggesting the recognition and misrecognition items are heterogeneous, yet the regression models treat individual items as coherent predictors. The low variance explained by the components, combined with the high missingness, means the current estimates rest on a fragile foundation.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This manuscript examines associations between self-reported recognition and misrecognition experiences in STEM contexts and two outcomes—STEM identity and STEM career aspirations—among undergraduate students at Minority Serving Institutions (N=1,134). Using stepwise multiple linear regression (for STEM identity) and logistic regression (for career aspirations), the authors find that specific recognition experiences (being called upon, teacher recommendation, receiving awards) significantly predict both outcomes, while misrecognition experiences do not. The study draws on discipline-based identity theory (Hazari et al., 2010) and uses a previously validated single-item STEM identity measure (Dou & Cian, 2022). The authors acknowledge that findings are preliminary due to high missingness (31% listwise deletion) and state their intention to apply multiple imputation in future work.","tokens_in":8229,"tokens_out":1386,"duration_ms":254076,"significance":"The study addresses a practically important question for STEM education research: which specific recognition experiences are most strongly associated with STEM identity and career aspirations among minoritized students. The use of a large, multi-MSI sample is a strength, and the focus on event-based recognition (rather than general perceptions) is a useful operationalization. The single-item STEM identity measure is drawn from published, validated work by one of the co-authors, which is standard practice. However, the manuscript is currently formatted as a conference presentation summary rather than a full journal article, and the analytic foundation has two load-bearing issues that must be addressed before the central claims can be considered well-supported.","major_comments":[{"comment":"§Findings, p. 8: The listwise deletion of 354 of 1,134 observations (31%) is the most serious concern. The authors acknowledge this and label findings as 'preliminary,' but the manuscript nonetheless draws substantive interpretive conclusions (e.g., that misrecognition is not associated with outcomes, that specific recognition experiences predict STEM identity). If missingness is related to either the predictors or outcomes (i.e., not MCAR), the reported coefficients and null results for misrecognition could be materially biased. The authors state they intend to use multiple imputation; this should be completed before the manuscript is submitted for journal review, as the current results are not publishable as stated. At minimum, a sensitivity analysis or missingness pattern diagnostic should be reported to assess the plausibility of MCAR.","section":null},{"comment":"§Analysis Procedures, p. 7: The PCA produces two components explaining only 43% of total variance (27% + 16%), yet the manuscript uses individual items (not component scores) in subsequent regressions. The relationship between the PCA and the regression models is unclear—if the PCA was conducted to address multicollinearity, the manuscript should explain how the PCA results informed the regression model specification. Currently, the PCA appears to be presented as justification but the regressions use the original items, leaving the analytical logic incomplete. The low variance explained also raises the question of whether the recognition items form a coherent construct, which should be discussed.","section":null},{"comment":"§Findings, p. 7: The logistic regression reports β = 0.47 (p = .002) for 'at least one recognition experience' predicting STEM career aspirations, with 1.59 higher odds. However, the manuscript also reports that individual recognition items (being invited to competitions, being accepted into STEM programs) were not significant. The construction of the 'at least one' variable and its relationship to the item-level results needs clarification—was this a composite variable, and if so, how was it created? The discrepancy between the composite and item-level findings should be explicitly discussed, as it bears on the interpretation of which specific experiences matter.","section":null}],"minor_comments":[{"comment":"The manuscript is formatted as a conference presentation summary (NARST 2025) rather than a full journal article. The level of methodological detail (e.g., stepwise regression entry/removal criteria, PCA rotation method, assumption checks) is insufficient for journal publication. The manuscript should be expanded to include full methodological transparency.","section":null},{"comment":"§Data sources, p. 4: The demographic breakdown notes that respondents could identify with multiple racial/ethnic categories, but the percentages (48% Black, 34% Hispanic, 31% white) sum to >100% without explicit clarification that this is expected due to multiracial identification. This should be noted at the point of presentation.","section":null},{"comment":"§Survey Items, p. 5: The recognition and misrecognition items are described but the response format (binary? Likert?) is not clearly specified for all items. The STEM identity item is described as 5-point Likert, but the recognition/misrecognition items appear to be binary ('whether they were...'). Clarifying the measurement level of each predictor would improve clarity.","section":null},{"comment":"§Analysis Procedures, p. 7: The stepwise regression approach is mentioned but the entry/removal criteria (e.g., p-to-enter, p-to-remove) are not specified. Given that stepwise methods are sensitive to these thresholds, they should be reported.","section":null},{"comment":"§Findings, p. 8: The linear regression reports R² = 0.1342, which is modest. The manuscript should discuss this in terms of practical significance and the large proportion of unexplained variance.","section":null},{"comment":"§Findings, p. 8: The statement that demographic variables 'were not found to be significantly associated with STEM career aspirations' is reported without the corresponding statistics. These should be reported in a table or in-text.","section":null},{"comment":"The retrospective self-report design—college students recalling K-12 experiences—is a limitation that should be explicitly acknowledged. While the claims are associational (not causal), the possibility of recall bias (e.g., students with high STEM identity over-reporting past recognition) should be discussed as a threat to validity.","section":null},{"comment":"Figure 1 (PCA biplot) is referenced but not included in the manuscript text provided. This should be included with clear labeling of component loadings in the final version.","section":null}],"recommendation":"major_revision","confidential_remarks":"This appears to be a conference presentation abstract/summary (NARST 2025) rather than a full manuscript prepared for journal submission. The authors themselves label findings as 'preliminary' and state they plan to conduct multiple imputation. I would recommend that the editor consider whether this submission is intended for a journal article or whether it was submitted in error. If it is intended as a journal submission, the authors should complete the planned analyses (multiple imputation) and expand the manuscript to full journal length with complete methodological reporting before re-submission. The underlying research design and research questions are reasonable and the sample is valuable, but the current manuscript is not yet at a stage suitable for journal review."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The referee raises three major concerns: (1) the 31% listwise deletion and associated risk of bias, (2) the unclear relationship between the PCA and the regression models, and (3) the construction and interpretation of the 'at least one recognition experience' composite variable. We agree with all three points and will revise the manuscript accordingly. Specifically, we will complete multiple imputation and report missingness diagnostics, clarify the role of the PCA, and explicitly describe the composite variable construction and discuss the discrepancy with item-level results. We note that the current manuscript is formatted as a NARST conference presentation summary, not a full journal article; the revisions described here will be implemented in the expanded journal-length version.","responses":[{"response":"We agree completely. The referee is correct that 31% listwise deletion is substantial and that the current results are not publishable as stated without further analysis. We labeled findings as 'preliminary' in the conference summary precisely because we recognized this limitation, but we concur that drawing substantive interpretive conclusions (including the null results for misrecognition) is premature given the missingness concern. In the revised manuscript, we will: (1) conduct and report Little's MCAR test or an equivalent missingness pattern diagnostic, (2) implement multiple imputation using chained equations (mice) with an appropriate number of imputed datasets, (3) re-estimate all regression models on the imputed data, and (4) report sensitivity analyses comparing results under listwise deletion versus multiple imputation. If the null results for misrecognition do not replicate under imputation, we will revise our conclusions accordingly. We will also temper all interpretive language until the imputation results are in hand.","revision_made":"yes","referee_comment":"The listwise deletion of 354 of 1,134 observations (31%) is the most serious concern. If missingness is not MCAR, the reported coefficients and null results could be materially biased. Multiple imputation should be completed before journal submission; at minimum, a sensitivity analysis or missingness pattern diagnostic should be reported."},{"response":"The referee is correct that the analytical logic is currently incomplete. The PCA was initially conducted as a diagnostic for multicollinearity among the recognition items, not as a basis for constructing component scores for the regressions. However, the manuscript does not make this logic explicit, and the reader is left uncertain about why the PCA is presented and how it informed the regression model specification. In the revision, we will: (1) clarify that the PCA served as a multicollinearity diagnostic rather than a variable-reduction step, (2) report the VIF statistics or condition indices that directly informed the multicollinearity assessment, and (3) discuss the low variance explained (43%) and what it implies about whether the recognition items form a coherent construct. If the PCA does not serve a clear purpose in the analytic pipeline, we will either remove it or restructure the analysis to use component scores consistently. We will also address whether the recognition and misrecognition items should be treated as separate constructs given the component structure.","revision_made":"yes","referee_comment":"The PCA produces two components explaining only 43% of total variance, yet the manuscript uses individual items (not component scores) in subsequent regressions. The relationship between the PCA and the regression models is unclear. The low variance explained also raises the question of whether the recognition items form a coherent construct."},{"response":"The referee raises a valid point. The 'at least one recognition experience' variable was created as a binary composite indicating whether a respondent reported any of the five recognition experiences (called upon, teacher recommendation, invited to competitions, accepted into STEM programs, received awards). The manuscript does not clearly describe this construction, nor does it explain why the composite is significant when two of the five constituent items are not. This discrepancy is important for interpretation: it could indicate that the composite captures a general 'recognition exposure' effect driven by the three significant items, or that the combination of experiences matters beyond any single item. In the revision, we will: (1) explicitly describe how the composite variable was constructed, (2) present both item-level and composite-level results in a single table for direct comparison, and (3) discuss the discrepancy, including whether the composite result is driven primarily by the three significant items and what this implies for the practical interpretation of which specific experiences matter. We will also consider whether the composite approach is appropriate given the item-level heterogeneity, or whether reporting only item-level results would be more informative.","revision_made":"yes","referee_comment":"The logistic regression reports β = 0.47 for 'at least one recognition experience' predicting STEM career aspirations, but individual recognition items (being invited to competitions, being accepted into STEM programs) were not significant. The construction of the 'at least one' variable and its relationship to the item-level results needs clarification, and the discrepancy should be explicitly discussed."}],"tokens_in":7268,"tokens_out":1080,"duration_ms":123447,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Here's the short version: this is a NARST conference paper reporting that specific recognition experiences (being called on, recommended for advanced classes, receiving awards) predict STEM identity and career aspirations in an MSI sample, while misrecognition experiences do not. The null on misrecognition is the most interesting finding because it cuts against prior work (Chachashvili-Bolotin et al., 2016; Simpson & Maltese, 2016). That contrast is a genuine new data point worth disseminating. The authors deserve credit for flagging their own limitations — they explicitly label findings as preliminary and plan multiple imputation. That honesty matters and is appropriately placed. The framework (Hazari et al., 2010; Dou & Cian, 2022) is well-established and the survey items are grounded in prior literature with expert review. The regression diagnostics are standard and the authors checked assumptions. Now the soft spots. The 31% listwise deletion is the load-bearing concern and the stress-test note is right to flag it. The authors acknowledge it, but the problem is real: if students with fewer recognition experiences skipped those items (plausible — you skip what you don't have), the analytic sample over-represents recognized students and inflates the associations. The null on misrecognition is equally vulnerable — differential missingness could produce a false null. Until multiple imputation is done, every coefficient and p-value here is provisional. The PCA is a secondary concern. Two components explaining 43% of variance is low, suggesting the recognition and misrecognition items are heterogeneous. The authors then shift to using individual items in the regressions rather than the components, which is actually defensible given the low variance explained — but the paper is unclear about when it's using components versus individual items, and that muddies the analysis. The retrospective self-report concern (college students recalling K-12 experiences) is real but standard for this kind of survey work. It's a limitation worth noting, not a disqualifier. The R² of 0.13 is low but not unusual for this kind of social science regression with attitudinal outcomes. The reader's scores are about right. Significance is modest — this confirms and extends existing frameworks rather than breaking new theoretical ground. The circularity concern is minimal; using a previously validated measure from a co-author's prior work is standard. Who is this for? STEM education researchers working on identity formation in minoritized populations, particularly people designing recognition-based interventions. It's a useful preliminary data point. My recommendation: this is a conference paper that the authors themselves call preliminary. It does not deserve a full journal-style peer review in its current form. The multiple imputation needs to happen first. If the results hold under MI, then it's worth a serious referee. Right now, it's an honest but incomplete analysis.","headline":"Conference paper with a genuinely interesting null result on misrecognition, but the analysis is preliminary — 31% listwise deletion and low PCA variance explained mean the current estimates are fragile.","tokens_in":8260,"tokens_out":662,"would_cite":false,"duration_ms":161798,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["01.40.Fk","01.40.J-"],"model":"glm-5.2","headline":"Being called on in class predicts STEM career hopes","keywords":[],"falsifier":"If a longitudinal study tracking recognition events in real time (rather than retrospectively) found no association between these specific events and later STEM identity or career aspirations, the core claim would be undermined.","tokens_in":7469,"feed_emoji":"🎓","tokens_out":623,"duration_ms":141866,"temperature":0.7,"pith_summary":"The paper examines whether specific moments of recognition by teachers and other adults during K-12 STEM education predict whether minoritized college students come to see themselves as STEM people and aspire to STEM careers. Surveying 1,134 students at Minority Serving Institutions, the authors find that three concrete experiences—being called on to answer questions in class, being recommended for an advanced STEM course, and receiving awards in STEM activities—significantly predict both STEM identity and career aspirations. Notably, misrecognition experiences (such as being passed over for recommendation or having negative classroom experiences) showed no significant association with either outcome. The authors argue that these findings suggest interventions should focus on engineering specific positive recognition moments rather than merely avoiding negative ones.","feed_headline":"","feed_subtitle":"","key_machinery":"The study uses principal components analysis to separate recognition experiences from misrecognition experiences into distinct factors, then applies logistic regression (for career aspiration) and multiple linear regression (for STEM identity) to test each factor's predictive power. The central analytical object is the distinction between event-based recognition and event-based misrecognition as independent predictors.","core_discovery":"The central finding is a dissociation between recognition and misrecognition: specific positive recognition events (being called on, recommended, awarded) predict STEM identity and career aspirations among minoritized students, while misrecognition events do not significantly predict either outcome. This challenges the assumption that negative experiences function as symmetrical impediments to positive ones. The study also finds that not all positive recognition is equal—being invited to competitions and being accepted into selective programs did not reach significance, whereas the three teacher-mediated classroom experiences did.","pith_inferences":[],"forward_implications":["Teacher training programs could be evaluated on whether they increase the frequency of specific recognition behaviors (calling on students, recommending for advanced courses) rather than generic encouragement","The asymmetry between recognition and misrecognition effects suggests that reducing negative experiences alone may be insufficient to boost STEM participation—active positive intervention is needed","The finding that competition invitations and selective program acceptance were not significant predictors could redirect intervention design away from competitive structures toward everyday classroom interactions","If the three significant recognition events are causal rather than merely correlational, they represent low-cost, scalable intervention targets that require no new programs or curricula"],"fun_headline_variants":["Teacher recognition predicts STEM identity; misrecognition shows no symmetrical effect","Positive STEM recognition events predict identity but misrecognition does not predict eith","Being called on and awarded predicts STEM identity more than selective program acceptance","Recognition and misrecognition are not symmetrical forces in minoritized STEM identity for","Teacher-mediated recognition predicts STEM identity; competition invitations did not reach"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The study asks college students to recall whether they were 'often called upon,' 'recommended by teachers,' or 'received awards' during their K-12 years, and treats these recollections as accurate records of past experiences. If students who already strongly identify as STEM people systematically remember more recognition than they actually received, the reported associations could reflect memory bias rather than genuine links between past events and current outcomes.","fun_headline_variants_meta":{"raw":{"variants":["Teacher recognition predicts STEM identity; misrecognition shows no symmetrical effect","Positive STEM recognition events predict identity but misrecognition does not predict either outcome","Being called on and awarded predicts STEM identity more than selective program acceptance","Recognition and misrecognition are not symmetrical forces in minoritized STEM identity formation","Teacher-mediated recognition predicts STEM identity; competition invitations did not reach significance"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":515,"prompt_tokens":419,"completion_tokens":96,"prompt_tokens_details":null},"tokens_in":419,"tokens_out":96,"duration_ms":49630,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T18:12:09.263035+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a longitudinal study tracking recognition events in real time (rather than retrospectively) found no association between these specific events and later STEM identity or career aspirations, the core claim would be undermined.","supporting_citations":[],"review_version":1}