{"id":"557847bf-0e97-4d33-b579-16edfd30f261","arxiv_id":"2605.27239","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Temporal simultaneity of annotations is the dominant predictor of inter-annotator agreement in a Setswana sentiment corpus, with kappa falling from 0.98 to 0.65 as time separation increases.","lead":"The paper analyzes a Setswana sentiment dataset and reports that inter-annotator agreement drops sharply over time, with temporal simultaneity as the strongest predictor. This observation could inform how future annotation campaigns for low-resource languages are scheduled to preserve quality.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Temporal simultaneity bins are likely collinear with sequential batch assignment, so observed κ gaps may reflect batch confounds rather than time per se.","rationale":"Reader correctly isolates the batch/time confound as the weakest link; the full-text analyses would need to demonstrate that the temporal result is robust to batch stratification for the dominant-predictor claim to hold. No other internal inconsistency appears from the reported numbers.","tokens_in":1762,"tokens_out":323,"duration_ms":23047,"concrete_test":"Re-estimate the temporal predictor after adding batch ID as a fixed effect (or by computing κ separately within each batch for the same time bins); if the <1 min vs >1 day κ gap falls below 0.15 or loses statistical significance, the simultaneity claim does not survive batch control.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the reported κ=0.98 (<1 min) vs κ=0.65 (>1 day) difference isolates a genuine simultaneity effect. Because the eight batches are sequential and per-batch κ already declines >32 points, any tweet pair separated by >1 day is disproportionately likely to straddle batches. Without an explicit control (batch fixed effects in the predictor model, within-batch temporal stratification, or regression of κ on time | batch), the simultaneity variable proxies unmeasured batch-level shifts in label distribution, annotator drift, or tweet difficulty. The abstract states annotation speed and tweet features show no association, but does not report the analogous check against batch identity.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a Setswana sentiment dataset of 3,565 tweets annotated by three native-speaker annotators across eight sequential batches. It reports an aggregate Randolph's free-marginal kappa of 0.76 that declines by more than 32 points across batches. Through six analyses the authors conclude that temporal simultaneity is the dominant predictor of per-tweet kappa (0.98 for labels within one minute versus 0.65 for labels more than one day apart), while annotation speed and tweet-level linguistic features show no association. The work also benchmarks multilingual encoders and proprietary models on three-class sentiment classification and releases the dataset, per-annotation timestamps, and analysis code.","tokens_in":1923,"tokens_out":498,"duration_ms":25627,"significance":"If the reported temporal-simultaneity effect survives explicit controls for batch identity, the finding would be a concrete, actionable contribution to annotation-quality management in long-running campaigns, especially for low-resource languages. The public release of timestamps and reproducible code is a clear strength that enables independent verification and extension.","major_comments":[{"comment":"Abstract: the central claim that temporal simultaneity is the dominant predictor of κ (0.98 vs. 0.65) is load-bearing. Because the eight batches are sequential and per-batch κ already declines >32 points, tweet pairs separated by >1 day are disproportionately likely to straddle batches. The manuscript does not report batch fixed effects, within-batch temporal stratification, or a regression of κ on time conditional on batch; without such a control the simultaneity variable may simply proxy unmeasured batch-level shifts in label distribution, annotator drift, or tweet difficulty.","section":"Abstract"},{"comment":"Abstract (six targeted analyses): the claim that annotation speed and tweet-level linguistic features show “no meaningful association” with κ is presented without the analogous check against batch identity. A direct comparison of the strength of the simultaneity predictor versus a batch-identity predictor is required to establish dominance.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: the model name “GPT-5” should be clarified (current model family or placeholder) and the exact few-shot and fine-tuning protocols should be summarized with the number of shots and training examples.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for identifying the need to explicitly control for batch identity when assessing the dominance of temporal simultaneity. We agree that the current analyses require these controls to strengthen the central claims and will revise the manuscript accordingly.","responses":[{"response":"We agree that the manuscript does not currently report batch fixed effects, within-batch stratification, or a conditional regression of κ on time. In revision we will add a linear regression of per-tweet κ on temporal separation that includes batch identity as a fixed effect, together with within-batch temporal stratification. These additions will test whether the simultaneity effect remains after accounting for batch-level variation.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that temporal simultaneity is the dominant predictor of κ (0.98 vs. 0.65) is load-bearing. Because the eight batches are sequential and per-batch κ already declines >32 points across batches. The manuscript does not report batch fixed effects, within-batch temporal stratification, or a regression of κ on time conditional on batch; without such a control the simultaneity variable may simply proxy unmeasured batch-level shifts in label distribution, annotator drift, or tweet difficulty."},{"response":"We acknowledge that the manuscript lacks a direct comparison of predictor strength between temporal simultaneity and batch identity. We will add this comparison (via regression coefficients or incremental R²) and will also re-evaluate the null associations for annotation speed and linguistic features after including batch fixed effects, to confirm that the reported dominance holds under these controls.","revision_made":"yes","referee_comment":"[Abstract] Abstract (six targeted analyses): the claim that annotation speed and tweet-level linguistic features show “no meaningful association” with κ is presented without the analogous check against batch identity. A direct comparison of the strength of the simultaneity predictor versus a batch-identity predictor is required to establish dominance."}],"tokens_in":1465,"tokens_out":422,"duration_ms":29599,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper reports a strong link between how close in time two annotators label the same tweet and their agreement level, with near-perfect kappa inside one minute falling to 0.65 across a day, and they position this as the dominant factor over speed or tweet features. They also release a new Setswana sentiment set with timestamps.\n\nWhat stands out as useful is the data release itself plus the per-annotation timestamps and analysis code. That makes the work reproducible for anyone auditing quality in similar low-resource settings. They run multiple checks on label confusion at the negative-neutral boundary and on run-length drift in two annotators, and the model benchmarks show clear gains from fine-tuning. Those pieces are concrete and worth having.\n\nThe soft spot is the batch structure. Eight batches run in sequence with per-batch kappa already declining more than 32 points, so tweets separated by more than a day are disproportionately likely to come from different batches. The abstract does not report a control for batch identity in the predictor model or a within-batch time stratification, even though it checks speed and linguistic features. Without that, the simultaneity variable may simply track unmeasured batch shifts in difficulty or annotator state rather than time per se. The stress-test note on collinearity looks like it lands.\n\nThis is for people building annotation pipelines for African-language or other low-resource data, especially if they care about practical scheduling. A reader focused on dataset quality gets the empirical observation and the released resources. It is coherent on its own terms and shows honest engagement with the literature on agreement decay, so it deserves a serious referee even if the causal isolation needs tightening.","headline":"Temporal simultaneity correlates with kappa here but likely proxies batch effects since batches are sequential and kappa already drops across them.","tokens_in":2427,"tokens_out":405,"would_cite":false,"duration_ms":24060,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The time gap between when annotators label the same tweet is the strongest predictor of their agreement.","keywords":["annotation quality","inter-annotator agreement","temporal simultaneity","sentiment analysis","Setswana","kappa","low-resource NLP","data quality"],"falsifier":"Re-labeling the same tweets while deliberately randomizing the time intervals between annotators and finding no corresponding change in kappa values would falsify the central claim.","tokens_in":2701,"feed_emoji":"⏱","tokens_out":653,"duration_ms":29558,"temperature":0.7,"pith_summary":"The paper examines a Setswana sentiment dataset of 3,565 tweets labeled by three annotators over eight batches and finds that overall Randolph's kappa is 0.76 yet per-batch agreement falls more than 32 points. Six analyses show the decline is driven mainly by temporal simultaneity rather than linguistic features or speed: tweets labeled within one minute reach kappa 0.98 while those labeled more than a day apart reach only 0.65. A sympathetic reader cares because long-running annotation campaigns for low-resource languages risk quality erosion unless time intervals are controlled. The work also identifies negative-neutral label confusion and run-length drift in some annotators before benchmarking models.","feed_headline":"Time between annotations predicts agreement quality","feed_subtitle":"Tweets labeled within one minute reach 0.98 kappa while those days apart fall to 0.65 in a Setswana sentiment dataset.","key_machinery":"temporal simultaneity, defined as the elapsed time between different annotators labeling the identical tweet","core_discovery":"The dominant predictor of inter-annotator agreement is temporal simultaneity: tweets labeled within one minute achieve κ = 0.98, while those labeled more than a day apart reach only κ = 0.65. Annotation speed and tweet-level linguistic features show no meaningful association with κ. Per-batch kappa declines across the task, label confusion concentrates on the negative/neutral boundary, and two annotators exhibit run-length drift.","pith_inferences":["Releasing per-annotation timestamps enables future datasets to audit for similar temporal quality loss.","High-agreement subsets filtered by short time gaps could serve as cleaner training data for models.","The pattern may appear in other low-resource annotation projects that span weeks or months.","Protocols that enforce same-day labeling could reduce the need for post-hoc quality filtering."],"forward_implications":["Annotation campaigns should schedule simultaneous labeling to sustain high agreement.","Label confusion is concentrated at the negative/neutral boundary.","Some annotators exhibit run-length drift consistent with autopilot labeling.","Fine-tuning yields 29 to 43 macro-F1 gains over pretrained baselines on the three-class task.","GPT-5 few-shot reaches the highest performance at 62.2 macro-F1."],"fun_headline_variants":["Annotation timing predicts sentiment label kappa","Simultaneous labels achieve 0.98 kappa","Day gaps between labels drop kappa to 0.65","Temporal simultaneity predicts annotation agreement"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The measured time gaps between annotations reflect a genuine effect on consistency rather than unmeasured differences in tweet difficulty or label distributions across batches.","fun_headline_variants_meta":{"raw":{"variants":["Annotation timing predicts sentiment label kappa","Simultaneous labels achieve 0.98 kappa","Day gaps between labels drop kappa to 0.65","Temporal simultaneity predicts annotation agreement"]},"model":"grok-4.3","cost_usd":0.006754,"raw_usage":{"total_tokens":3163,"prompt_tokens":708,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":67537000,"prompt_tokens_details":{"text_tokens":708,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2402,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":708,"tokens_out":53,"duration_ms":29466,"temperature":1.0,"reasoning_tokens":2402,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:16:18.506462+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-labeling the same tweets while deliberately randomizing the time intervals between annotators and finding no corresponding change in kappa values would falsify the central claim.","supporting_citations":[],"review_version":1}