{"id":"f7fc8c42-1223-4e62-a10d-2718f483346b","arxiv_id":"2602.02208","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AgriHubi, a Finnish-language agricultural RAG system built on PORO models, showed improved user ratings (top scores from 3% to 21%) across two rounds of testing.","lead":"This paper describes AgriHubi, a Finnish-language agricultural question-answering system that combines government documents with PORO language models. It reports that user satisfaction scores improved across eight design iterations, with top ratings rising from 3% to 21%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of 'clear gains' is undermined by the uncontrolled before/after design: system changes and question-set differences are confounded, and the rating shift may reflect task difficulty or response style rather than system improvement.","rationale":"The reader identified the comparability of the two user studies as the weakest assumption, and this is precisely the load-bearing concern. The paper's own limitation statement in Section V-C acknowledges differences in question types, and the system changed in multiple dimensions simultaneously (model, retrieval, chunking, answer length). The reader's verdict of CONDITIONAL is appropriate: the evidence is suggestive but not conclusive. I agree with the reader's assessment and do not see a reason to change the verdict. My additional point about the lack of direct measurement of 'completeness' and 'linguistic accuracy' reinforces, rather than replaces, the comparability concern. The proposed concrete test — expert-rated question difficulty stratification or a crossover study — would settle whether the observed rating shifts are attributable to the system changes.","tokens_in":9100,"tokens_out":3835,"duration_ms":37882,"concrete_test":"Obtain the complete set of questions used in both evaluation rounds. Have an independent panel of Finnish-speaking agricultural experts rate each question for difficulty (e.g., number of reasoning steps, domain specificity, ambiguity) without knowing which round it came from. Then stratify the user ratings by question difficulty. If August questions are rated systematically easier or more answerable, the improvement in ratings may be an artifact of question selection rather than system improvement. A stronger test would be a within-subject crossover: have the same users answer the same questions with both the April and August system versions in randomized order; if the rating difference shrinks or reverses, the original before/after comparison is confounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core claim is that AgriHubi improved 'in answer completeness, linguistic accuracy, and perceived reliability' based on two user studies (April vs August, Table I). The load-bearing assumption is that the two evaluation rounds are comparable apart from the system version. This assumption is not met. Between rounds, the system changed simultaneously in generative model (PORO-34B vs PORO-2-70B/8B, Section IV-C), retrieval strategy and chunking (Iterations 5 and 8), and maximum answer length (700 to 2000 tokens). The paper itself notes in Section V-C that 'differences in question types between evaluation rounds may have introduced minor scoring bias' — but this is not minor if August questions were easier or more answerable, because no control for question difficulty is described. The rating shift (low ratings 46%→38%, top ratings 3%→21%) could arise from easier questions, longer (not necessarily more correct) answers, or different user expectations — especially since the evaluation was not blinded and users knew they were testing an 'updated' system. Furthermore, the claims of improved 'completeness' and 'linguistic accuracy' are not directly measured; users gave a single holistic 1–5 rating, and no per-criterion scoring or independent analysis of answer text is reported. Therefore, the central claim is underdetermined: the comparison does not isolate the system's quality as the cause of the observed improvement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AgriHubi, a retrieval-augmented generation (RAG) system for Finnish-language agricultural decision support, built on Finnish agricultural documents and open PORO-family language models. The system was developed over eight iterations and evaluated in two user studies (April 2025, 67 ratings; August 2025, 47 ratings). The authors report that low ratings (1–2) decreased from 46% to 38% and top ratings (5) increased from 3% to 21% (Table I), and they interpret this as clear gains in answer completeness, linguistic accuracy, and perceived reliability. The paper also discusses latency/quality trade-offs and provides qualitative observations from user feedback.","tokens_in":9440,"tokens_out":2913,"duration_ms":30707,"significance":"If the central claim were established, AgriHubi would be a useful case study for domain-specific RAG in a low-resource language: it combines an open-source Finnish-capable model family with localized documents, integrates a user-feedback loop, and reports real-world deployment lessons. The paper is transparent about many architectural choices and limitations, and the system description is sufficiently detailed to be reimplementable. However, the headline claim of 'clear gains' rests on an uncontrolled before/after comparison with small samples, no statistical testing, and no direct measurement of the named quality dimensions. As it stands, the paper is better characterized as a system description and observational report than as evidence that the deployed system improved in the measured qualities. These gaps are load-bearing for the paper's main conclusion.","major_comments":[{"comment":"The central claim of improvement is underdetermined by the April-to-August comparison. Between the two rounds the generative model changed (PORO-34B to PORO-2-70B/8B, Section IV-C), retrieval strategy and chunking changed (Iterations 5 and 8), and maximum answer length increased from 700 to 2000 tokens (Section IV-C). The paper itself acknowledges that 'differences in question types between evaluation rounds may have introduced minor scoring bias' (Section V-C). With no control for question difficulty, user composition, or rating behavior, the observed shift in Table I cannot be attributed to system quality. I recommend either re-analyzing the data with a matched question set or paired user ratings, or substantially softening the causal language throughout.","section":"Section V-B / Table I"},{"comment":"The abstract and Section I claim gains specifically in 'answer completeness, linguistic accuracy, and perceived reliability,' but the only quantitative instrument reported is a single holistic 1–5 Likert rating. There is no rubric, no per-criterion scoring, and no independent evaluation of the answer text for completeness or linguistic accuracy. Therefore the named dimensions are not directly operationalized. The paper should either provide per-criterion evaluation data or restrict the claims to what the rating measures, i.e., overall user-perceived answer quality.","section":"Section V-A / Table I"},{"comment":"The quantitative evidence lacks inferential statistics. With 67 and 47 ratings, the changes in proportions (46% to 38% low; 3% to 21% top) may be within sampling variability. No significance tests, confidence intervals, or effect sizes are reported. The authors acknowledge in Section V-C that the sample sizes were 'insufficient for fine-grained statistical analysis,' but this admission appears only in Limitations; the Results section presents the differences as definitive. At minimum, report exact tests (e.g., Mann-Whitney U, chi-square, or bootstrapped proportion differences) and interpret the results accordingly.","section":"Section V-B / Section V-C"},{"comment":"The qualitative feedback is summarized without a systematic protocol. The paper does not describe how many users provided written comments, how the comments were collected, or how themes like 'clearer and more reliable' were derived. No representative quotations or inter-rater reliability are given. In addition, users were not blinded to the system version, so expectancy effects could contribute to the more positive August feedback. The qualitative evidence should be presented as illustrative rather than as confirmation of improvement.","section":"Section V-D / Section VI"}],"minor_comments":[{"comment":"The phrase 'clear gains' is stronger than the evidence supports; consider 'reported improvements' or 'observed increases' in the abstract and conclusion.","section":"Abstract / Section V-B"},{"comment":"The number of participants is not stated; only the number of responses (67 and 47) is given. If multiple ratings per user are included, the effective sample size and potential within-user correlation should be addressed.","section":"Section V-A"},{"comment":"The notation 'PORO-2-70B/8B' in the evaluation section is ambiguous; specify which model was used for each response, since Section V-A says 'in some cases PORO-2-8B' was used.","section":"Section IV-C"},{"comment":"The statement that question-type differences 'may have introduced minor scoring bias' should be justified. If the bias is assumed minor, explain why; otherwise the assumption undermines the comparison.","section":"Section V-C"},{"comment":"Figures 1 and 2 are referenced but not visible in the manuscript text; ensure the submitted version includes them and that they are legible.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable system description with an honest limitations section, but the evaluative claims outstrip the evidence. The confounded before/after design is the core problem. I would not advocate rejection because the system description and lessons learned have value, and the claims could be reframed as observational. However, the authors need to either supply stronger quantitative analysis (matched questions, statistical tests, per-criterion ratings) or explicitly demote the 'clear gains' claim to a preliminary observation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate case study, not a breakthrough. The genuinely new content is the specific Finnish-language agricultural RAG system, AgriHubi, built on PORO models, plus two rounds of user ratings and qualitative feedback. The paper is clearly written and unusually honest about its limitations.\n\nWhat it does well: the eight-iteration development story is concrete and believable. The choice of PORO models for Finnish is sensible, the retrieval pipeline is standard but competently assembled, and the reported latency/quality trade-off matches prior findings. The qualitative feedback aligns with the rating shift, which makes the narrative internally consistent.\n\nSoft spots: the central claim of 'clear gains' rests on a before/after comparison where the system changed in several ways at once — model, retrieval strategy, chunking, and maximum answer length — and the question sets were not identical across rounds. The paper acknowledges in Section V-C that differences in question types 'may have introduced minor scoring bias,' but 'minor' is not established. If August questions were easier or more answerable, the observed shift could reflect task difficulty rather than system improvement. The rating is a single holistic 1–5 scale, so the specific claims about completeness and linguistic accuracy are not directly measured; two raters could have different criteria. Sample sizes (67 and 47) are small, no significance tests or confidence intervals are reported, and the code and data are not public, so independent reproduction is impossible.\n\nThat said, the stress-test note is fair but not damning. The paper itself flags most of these issues in the limitations section, and the design is typical for an iterative engineering case study. The problem is that the abstract and conclusion use 'clear gains' without hedging. That wording should be tempered to something like 'reported improvements' or 'perceived gains,' and the authors should either provide artifacts or explicitly frame the result as a case study with no causal attribution.\n\nThe citation pattern is fine; the two self-cited prior works are background, not load-bearing. The paper is a useful descriptive contribution to an under-served area (low-resource language RAG). It deserves a serious referee — not because it is rigorous, but because the community needs more real-world deployments reported, and the evaluation weaknesses are fixable with revisions. I would not accept it as-is, but I would engage with it.","headline":"A readable, honest case study of Finnish agricultural RAG, but the central 'clear gains' claim is undercut by an uncontrolled before/after design and a holistic rating scale.","tokens_in":9919,"tokens_out":2796,"would_cite":false,"duration_ms":27090,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Finnish-language agricultural RAG system, grounded in local documents and refined through user feedback, shows clear gains in answer quality and user-rated reliability.","keywords":["retrieval-augmented generation","Finnish language","agricultural decision support","PORO models","user evaluation","low-resource language","iterative system design","LLM feedback loop"],"falsifier":"Run the August system and the April system on the same fixed set of Finnish agricultural questions with the same group of raters, switching only the model and pipeline; if the 46%→38% low-rating and 3%→21% top-rating shifts disappear, the reported gains are an artifact of the simultaneous changes.","tokens_in":9008,"feed_emoji":"🌾","tokens_out":5612,"duration_ms":45903,"temperature":0.7,"pith_summary":"This paper tries to establish that a retrieval-augmented generation (RAG) system built specifically for Finnish-language agriculture—AgriHubi—can move from prototype to usable decision-support tool through iterative, feedback-driven refinement. The authors claim that combining Finnish agricultural PDFs with open PORO-family language models and an explicit source-grounding mechanism improves answer completeness, linguistic accuracy, and perceived reliability. They support this with two user studies: between April and August 2025, the share of low user ratings (1–2) fell from 46% to 38%, and top ratings (5) rose from 3% to 21%. The paper argues that the largest gains came from retrieval and preprocessing improvements rather than model scaling alone.","feed_headline":"Finnish farm AI: top user ratings rise from 3% to 21%","feed_subtitle":"Eight feedback-driven rounds turn a Finnish farm AI prototype into a trusted tool.","key_machinery":"The central mechanism is the domain-adapted RAG pipeline: preprocessing Finnish agricultural PDFs with OCR, chunking with metadata, embedding via text-embedding-ada-002 into a FAISS index, retrieval of top-k chunks with L2-normalized cosine similarity, and prompt assembly through a language handler that detects Finnish, Swedish, and English. The PORO family models generate answers, and a SQLite database logs every query, retrieved passage, response, and rating—this feedback loop drives the eight iterations.","core_discovery":"AgriHubi integrates Finnish agricultural documents with PORO family models in a RAG pipeline connecting document store, FAISS retriever, generative model, and a chat interface with built-in five-point ratings. Over eight iterations, it evolved from a Llama-based prototype to a Finnish-optimized platform. The paper's central discovery: iterative refinement guided by user feedback—especially retrieval, chunking, and answer-length changes—produced measurable gains. In the April round (PORO-34B, 67 ratings), 46% of answers were rated 1–2 and 3% rated 5; in August (PORO-2-70B, 47 ratings), low ratings dropped to 38% and top ratings rose to 21%. The authors attribute gains mainly to system-level r","pith_inferences":["If the latency-trust link is causal, optimizing response time in shared-GPU deployments may raise perceived reliability even without further retrieval changes; a direct test would compare user ratings under artificially slowed versus full-speed responses.","The same feedback loop could transfer to other low-resource language domains with national documentation (e.g., legal or medical Finnish); running AgriHubi's eight-iteration process on a different language-document pair would test that transferability.","Because the paper changed model, retrieval, chunking, and answer length simultaneously, the individual contribution of each factor is unresolved; a factorial experiment would isolate them and could reveal which single change carries the largest share of the gain."],"forward_implications":["If correct, domain-specific RAG systems for low-resource languages can be built and improved by combining local documents with open multilingual models, without large English-centric fine-tuning.","Iterative, feedback-driven refinement of retrieval and preprocessing, rather than model scaling alone, can turn a prototype into a usable decision-support tool within months.","Larger models within the PORO family give more accurate and complete answers but increase latency, and latency directly shapes user trust in deployed systems.","Structured user studies with logged ratings and qualitative comments provide a workable evaluation method for tracking system improvement over iterations.","Human review remains necessary for catching subtle terminology and context errors that automated metrics miss in low-resource settings."],"fun_headline_variants":["Feedback loops lift Finnish farm AI top ratings from 3% to 21%","AgriHubi iterates to 21% top ratings in Finnish agricultural AI","User-driven refinement boosts Finnish farm AI reliability","From 3% to 21%: feedback-driven RAG improves Finnish farm AI","Domain-specific RAG system for Finnish agriculture gains trust"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The two user studies are comparable enough that the improvement in ratings is caused by the system changes rather than by differences in user groups, question difficulty, or rating behavior.","fun_headline_variants_meta":{"raw":{"variants":["Feedback loops lift Finnish farm AI top ratings from 3% to 21%","AgriHubi iterates to 21% top ratings in Finnish agricultural AI","User-driven refinement boosts Finnish farm AI reliability","From 3% to 21%: feedback-driven RAG improves Finnish farm AI","Domain-specific RAG system for Finnish agriculture gains trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000312,"raw_usage":{"total_tokens":1591,"prompt_tokens":700,"completion_tokens":891,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":798}},"tokens_in":444,"tokens_out":891,"duration_ms":7248,"temperature":1.0,"reasoning_tokens":798,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:24:16.999391+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the August system and the April system on the same fixed set of Finnish agricultural questions with the same group of raters, switching only the model and pipeline; if the 46%→38% low-rating and 3%→21% top-rating shifts disappear, the reported gains are an artifact of the simultaneous changes.","supporting_citations":[],"review_version":1}