{"id":"1fbf34f5-7823-423b-b1d7-616922fee690","arxiv_id":"2503.04732","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An ongoing Brazilian Informatics Olympiad training program reports descriptive gains in participation and qualifier numbers, with no control group or statistical analysis yet.","lead":"This paper reports an ongoing training program that prepares Brazilian high school students for the Informatics Olympiad (OBI) and shares first-round participation data from 2023 and 2024. It is an early experience report: the authors describe growth in registrations and phase-2 qualifiers, while cautioning that the data do not yet show that the training caused these gains.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim overreaches: Section 7's 'demonstram' is not supported because the only evidence is year-over-year growth in participation and qualifiers with voluntary enrollment, no control group, and no skill measure; average scores did not improve.","rationale":"I read the paper as a preliminary experience report whose intended contribution is to describe a training methodology and report first-cycle OBI outcomes. For the central claim to hold, the observed increases must be attributable to the training, and participation and qualification must be valid proxies for programming skill and computational thinking. The first condition is the least secure: Section 4.2 describes voluntary enrollment with no selection or randomization; Section 5 acknowledges the data 'podem sugerir' but that 'análise mais detalhada ainda será realizada'; Section 7 elevates this to 'demonstram.' The second condition is also strained because Section 5 reports average scores did not meaningfully improve and no independent skill assessment is reported. These are not accusations of dishonesty; the authors are appropriately cautious in the results section, and the DBR plan is a reasonable way forward. The concrete test—a difference-in-differences comparison against non-trained institutions—would directly test whether the observed growth is specific to the program. If it shows parallel trends, the conclusion must be softened to 'suggestive'; if it shows treatment-specific gains, the central claim gains support. Since the reader's CONDITIONAL verdict already captures this gap, no change to the verdict is needed.","tokens_in":8586,"tokens_out":3671,"duration_ms":34797,"concrete_test":"Obtain OBI participation and phase-2 qualification microdata for 2023 and 2024 for all IFTM campuses and for a matched set of non-trained federal or state schools in the same region, then run a difference-in-differences analysis comparing treated versus untreated schools before versus after the training. If non-trained schools show similar growth in participation or qualification rates, the training effect is not identified; if the growth is specific to treated campuses and robust to school fixed effects, the causal claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the causal inference in Section 7: 'os resultados preliminares demonstram que treinamento em programação é uma ferramenta relevante para o ensino de programação e promoção do pensamento computacional.' The observed facts—IFTM campuses participating growing from 4 to 6, P1 qualifiers rising by 8, and female P1 participation at Uberlândia Centro growing from 3 to 8—are compatible with the training having an effect, but they do not demonstrate it. Enrollment was voluntary (Section 4.2), so participants were self-selected and likely more motivated than non-participants. There is no comparison group, no pre-test, and no statistical test. Section 5 explicitly hedges: 'Estes dados podem sugerir um reflexo dos treinamentos oferecidos... no entanto, uma análise mais detalhada ainda será realizada.' The conclusion drops that hedge. Moreover, the proxy assumption—that participation and phase-2 qualification measure programming skill or computational thinking—is weakened by the paper's own finding that average OBI scores in 2024 were not significantly different from 2023 (Section 5, Figs. 7-8). The increase in qualifiers could reflect cohort composition, national OBI trends, or familiarity with the exam format rather than learning. The paper is an honest preliminary report, but the central claim as worded needs to be softened to 'consistent with' or 'suggestive of,' pending the planned DBR cycles and formal analysis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on an ongoing Design-Based Research project to develop a methodology for teaching programming to Brazilian middle and high school students through competitive programming training, centered on the Brazilian Informatics Olympiad (OBI). It describes the training setup (voluntary enrollment, presencial and remote modalities, Neps Academy and Beecrowd tools), presents descriptive statistics comparing 2023 and 2024 OBI participation across IFTM campuses, and concludes that preliminary results demonstrate the relevance of programming training for teaching programming and promoting computational thinking.","tokens_in":8789,"tokens_out":3382,"duration_ms":32900,"significance":"If the causal claim were supported, the paper would provide a useful, low-cost model for increasing OBI participation and potentially fostering computational thinking in Brazilian basic education. The paper's strengths include an honest acknowledgment that the research is preliminary, a clear description of the training methodology that can be replicated by others, and the explicit caveat in Section 5 that a more detailed analysis is still pending. The descriptive data on campus expansion, female participation, and P1 qualifiers are useful starting points. However, the central claim as worded in Section 7 is not supported by the evidence, which is entirely descriptive and lacks a control group or independent skill measure.","major_comments":[{"comment":"The claim that 'os resultados preliminares demonstram que treinamento em programação é uma ferramenta relevante para o ensino de programação e promoção do pensamento computacional' overreaches. The supporting evidence in Section 5 is year-over-year growth in the number of participating campuses (4 to 6), an increase of 8 P1 qualifiers, and an increase in female P1 participation at one campus (3 to 8). Because enrollment was voluntary (Section 4.2), there is no control group, no pre-test, and no statistical test; these observations are equally compatible with self-selection, national OBI trends, or school-level differences. The authors themselves hedge in Section 5 ('Estes dados podem sugerir um reflexo dos treinamentos oferecidos... no entanto, uma análise mais detalhada ainda será realizada'), but this hedge is dropped in the conclusion. The conclusion should be softened to 'consistent with' or 'suggestive of' and explicitly framed as preliminary pending the planned DBR cycles and formal analysis.","section":"Section 5, Figures 7-8"},{"comment":"The paper assumes that OBI participation and phase-2 qualification are valid proxies for programming skill and computational thinking, but this assumption is weakened by the paper's own finding that average OBI scores in 2024 were not significantly different from 2023. If learning had occurred, one would expect some improvement in scores, especially among returning or trained students. The increase in qualifiers could reflect larger cohort sizes, cohort composition, familiarity with the exam format, or favorable contest conditions rather than skill gains. The manuscript should either report a direct pre/post measure of skill (e.g., scores on training exercises or a separate test) or restrict the central claim to outcomes of participation and interest, which the data can more directly support.","section":"Section 5, Figures 7-8"},{"comment":"The quantitative analysis is limited to raw counts and percentages, with no confidence intervals, effect sizes, or significance tests. The observed increases involve small numbers (e.g., female P1 participation rising from 3 to 8), making it difficult to judge whether these changes are meaningful or stable. Furthermore, the paper does not report how many of the students who took the OBI exam were actually participants in the training program, so the link between the training and the observed outcomes is not established. The authors should provide a breakdown of trained versus untrained test-takers and, at minimum, report the uncertainty around the reported counts before drawing any causal or even associational conclusions.","section":"Section 5, Table 2 and Figures 2-6"}],"minor_comments":[{"comment":"The phrase 'em relação ao ano de 2024' appears to be a typo for 'em relação ao ano de 2023', since the sentence compares 2023 (10 female participants) with 2024 (14 female participants).","section":"Section 5, paragraph on female participation"},{"comment":"The text first says that 'O Campus Uberlândia Centro e Uberaba participaram da prova em 2023' and later says that 'É a primeira participação na OBI na história do Campus Uberlândia.' Clarify whether 'Campus Uberlândia' and 'Campus Uberlândia Centro' are distinct campuses, and if so, which one is which, because the current wording is confusing.","section":"Section 5, paragraph on campus expansion"},{"comment":"The captions for Figures 7 and 8 are identical to those of Figures 5 and 6 ('Participantes por nível - 2023' and '2024'), but the text states that Figures 7 and 8 show average scores. The figure captures should be corrected to reflect that they show performance means, not participant counts.","section":"Section 5, Figures 7 and 8"},{"comment":"The abstract and conclusion use the verb 'demonstra', while Section 5 explicitly hedges the interpretation. Align the abstract and conclusion with the more cautious language used in the results section, or provide the additional analysis needed to justify the stronger wording.","section":"Abstract and Section 7"},{"comment":"The text mentions 'um aumento de 8 alunos' for P1 qualifiers, but Table 2 itself is not described with the baseline and post values. State the absolute numbers (e.g., from N to N+8) in the text so the reader does not have to infer them from the table alone.","section":"Section 5, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short experience report that describes an ongoing project; its main technical weakness is the gap between the descriptive data and the strong causal language in the conclusion. This gap is fixable by softening the claims and adding caveats, so major revision rather than rejection is appropriate. The authors may also want to consider whether the paper's contribution is best framed as a descriptive case study rather than an evaluation of training effectiveness at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is an honest, clearly written experience report with new observational data from the 2024 OBI cycle at IFTM. The descriptive numbers are internally consistent, and the authors repeatedly flag the preliminary status. The problem is that the conclusion in Section 7 drops the hedge: 'demonstram' is too strong for evidence that has no control group, no pre-test, no causal identification, and no independent skill measure.\n\nWhat's actually new is modest but real: campus-level participation counts, gender breakdowns, level distribution, and phase-2 qualifiers for the 2024 cycle, which don't appear in the cited Menezes thesis. The training method itself is borrowed from that thesis and prior experience reports, so the novelty is in the new cohort and new campuses, not in the pedagogy.\n\nThe paper does several things well. It is transparent about methodology: voluntary enrollment, the tools used (Neps Academy, Beecrowd), the disruption from the IFTM strike, and the need for deeper analysis. The Section 5 sentence 'Estes dados podem sugerir... no entanto, uma análise mais detalhada ainda será realizada' is exactly the right level of caution. The planned design-based research cycles are a reasonable path to stronger evidence.\n\nThe soft spot is the central claim. The observed increase in participating campuses (from 4 to 6) and P1 qualifiers (up 8) is compatible with training effects, but voluntary enrollment means self-selection is a live confounder. The paper's own finding that average scores did not improve weakens the proxy assumption that participation and qualification measure programming skill. The conclusion should be softened to 'consistent with' or 'suggestive of,' pending the formal analysis. These are fixable with revision; the internal logic of the report holds.\n\nWho is this for? People running outreach training programs in computing education, especially in Brazil, will find the operational details useful. It is not a rigorous evaluation and should not be cited as evidence of effectiveness.\n\nRecommendation: send it to peer review. It is a legitimate experience report that needs a modest revision to align conclusions with evidence. Desk rejection would be too harsh, but a careful referee should push for softened language and a future comparison group.","headline":"Honest experience report with new 2024 OBI data, but the Section 7 conclusion overreaches: the evidence supports 'suggestive of' rather than 'demonstrates.'","tokens_in":9438,"tokens_out":1519,"would_cite":false,"duration_ms":14797,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Competitive-programming training can teach programming and promote computational thinking in high school.","keywords":["competitive programming","Brazilian Informatics Olympiad","computational thinking","high school education","programming training","online judge platforms","C++ study tracks","gender participation"],"falsifier":"Compare the same participation and qualification metrics in similar schools or campuses that did not receive the training over the same 2023-2024 period; if the increases match untrained schools, the program's causal role is unsupported. Alternatively, administer a validated pre/post computational-thinking test to trained students: if scores do not move while participation rises, the proxy claim that participation reflects skill growth would be contradicted.","tokens_in":8330,"feed_emoji":"💻","tokens_out":8552,"duration_ms":76703,"temperature":0.7,"pith_summary":"This paper reports early results from an ongoing project that trains Brazilian middle- and high-school students for the Brazilian Informatics Olympiad (OBI). The authors' central claim is that competitive-programming training is a relevant tool for teaching programming and promoting computational thinking. Their preliminary evidence is a set of year-over-year comparisons: the number of participating campuses at the target federal institute grew from four to six, phase-2 qualifiers in the P1 level rose by eight students, and female P1 participation on one campus grew from three to eight. They interpret these as positive signs that volunteer outreach plus structured training can broaden access to computing competitions, while stating that a deeper analysis is still to come.","feed_headline":"Volunteer coding training lifts OBI participation and qualifiers","feed_subtitle":"More campuses entered Brazil's Informatics Olympiad and more students reached phase 2 after volunteer training.","key_machinery":"The load-bearing object is a volunteer training cycle: outreach to students, synchronous in-person and remote classes built around C++ study tracks on an online learning platform, practice on an online judge system, and intensified simulations before the OBI. The mechanism that carries the claim is the contrast between OBI participation metrics in 2023 and 2024: campus count, per-level enrollment, female participation, and phase-2 qualifiers.","core_discovery":"The paper's central claim is that competitive-programming training is a relevant, workable tool for teaching programming and promoting computational thinking in Brazilian middle and high school. Its evidence is early and correlational: participation by campuses grew from four to six, P1 qualifiers for the second phase rose by eight, and female P1 enrollment on one campus grew from three to eight. The authors describe these as preliminary signs that volunteer training with online judges can both widen access to scientific competitions and support the computational-thinking goals of Brazil's national curriculum.","pith_inferences":["Because enrollment was voluntary and no control group was used, the observed year-over-year growth could also reflect preexisting interest or other school-level trends rather than the training itself; a comparison with untrained schools would separate these explanations.","The jump in female participation on the campus taught by a woman instructor suggests a testable hypothesis: visible female role models in competitive programming may lower the participation barrier, which later cycles could test by varying instructor gender across campuses.","Participation counts and qualification numbers are the only outcome measures so far; a validated pre/post computational-thinking test would tell whether the training builds skill or mainly attracts and mobilizes students who already have interest.","If the model proves causal in further cycles, it would offer a low-cost, volunteer-run blueprint for expanding computer science education in school systems with limited computing infrastructure."],"forward_implications":["If the preliminary results are representative, one volunteer training cycle can bring additional campuses into the OBI and increase the number of students reaching the second phase.","Extending the training to five more cycles through 2026 should produce stronger evidence on qualifiers, medals, and the methodology's generalizability.","The more than two-fold rise in female P1 participation on the campus with a female instructor suggests that outreach and instruction style may matter for gender inclusion, a direction the project's next cycles can test.","The authors' decision to keep training students who did not qualify implies that repeated exposure, not just one competition, is part of the proposed pathway to improved performance.","The lack of average score gains in 2024 is not treated by the authors as a failure; they attribute it to novice students and a strike that reduced attendance, framing participation growth as the current success metric."],"supporting_citations":[{"why":"Defines the OBI's structure, levels, and phases, which supply the outcome metrics the paper uses.","marker":"[SBC 2022]"},{"why":"Provides the training method that this paper's cycles are based on.","marker":"[Menezes 2024]"},{"why":"Identifies the online learning platform used in the training as the one preferred by OBI medalists.","marker":"[de Menezes et al. 2021]"},{"why":"Supplies the definition of computational thinking that the paper claims the training promotes.","marker":"[Wing 2006]"},{"why":"Gives the iterative design methodology used to refine the training across cycles.","marker":"[Herrington et al. 2007]"},{"why":"Frames competitive programming as a form of problem-based learning, connecting the training to the pedagogical approach.","marker":"[Laaksonen 2017]"},{"why":"Earlier experience report that structured training builds a culture of competitive programming, which this paper extends to high school.","marker":"[Sousa et al. 2021]"}],"fun_headline_variants":["Coding training lifts OBI participation and qualifiers","Training grows OBI entrants and second-phase qualifiers","Volunteer training widens OBI access, early data shows","OBI participation climbs after competitive programming training","Preliminary signs: training boosts OBI participation and qualifiers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed year-over-year increases in OBI participation and qualifier counts are caused by the training program, rather than by unrelated trends, school-level differences, or changes in who chose to enroll.","fun_headline_variants_meta":{"raw":{"variants":["Coding training lifts OBI participation and qualifiers","Training grows OBI entrants and second-phase qualifiers","Volunteer training widens OBI access, early data shows","OBI participation climbs after competitive programming training","Preliminary signs: training boosts OBI participation and qualifiers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000584,"raw_usage":{"total_tokens":2641,"prompt_tokens":735,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":351,"completion_tokens_details":{"reasoning_tokens":1826}},"tokens_in":351,"tokens_out":1906,"duration_ms":13368,"temperature":1.0,"reasoning_tokens":1826,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T19:32:11.403818+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the same participation and qualification metrics in similar schools or campuses that did not receive the training over the same 2023-2024 period; if the increases match untrained schools, the program's causal role is unsupported. Alternatively, administer a validated pre/post computational-thinking test to trained students: if scores do not move while participation rises, the proxy claim that participation reflects skill growth would be contradicted.","supporting_citations":[{"cited_title":"Olimpíada brasileira de informática","cited_arxiv_id":null,"evidence_quote":"Defines the OBI's structure, levels, and phases, which supply the outcome metrics the paper uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the training method that this paper's cycles are based on."},{"cited_title":"R., de Souza Pereira, J","cited_arxiv_id":null,"evidence_quote":"Identifies the online learning platform used in the training as the one preferred by OBI medalists."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of computational thinking that the paper claims the training promotes."},{"cited_title":"C., and Oliver, R","cited_arxiv_id":null,"evidence_quote":"Gives the iterative design methodology used to refine the training across cycles."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames competitive programming as a form of problem-based learning, connecting the training to the pedagogical approach."},{"cited_title":"R., Silva, G., Lima, V., Tavares, W., and Bezerra, C","cited_arxiv_id":null,"evidence_quote":"Earlier experience report that structured training builds a culture of competitive programming, which this paper extends to high school."}],"review_version":1}