{"id":"142d647c-ac0c-4014-ad58-34002e0a3b0e","arxiv_id":"2501.07487","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative review of sustainable AI challenges and solutions across data acquisition, data processing, and model training, with RISC-V highlighted as a hardware opportunity.","lead":"This paper surveys how AI can be made more sustainable by improving data collection, processing, and hardware/software design. It is a high-level review that adds no new results and may contain several fabricated citations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The RISC-V-enabler recommendation rests on unverified performance figures; §4.2.1's 2–3× Xuantie claim and §4.1.1's 1280 MWh figure lack primary sources, so the central claim is not yet supported.","rationale":"The reader's verdict of UNVERDICTED is appropriate: this is a review/position paper with no original experiments, and the quantitative support for its central recommendation is less secure than the prose suggests. I read the strongest claim charitably as a call to treat data acquisition, data processing, and model training/inference as connected levers for reducing AI's environmental footprint, with open hardware such as RISC-V as a concrete enabler. The load-bearing precondition is that the cited technologies actually deliver the claimed efficiency gains. That precondition is exactly where the evidence is weakest. Section 4.2.1's 2–3× Xuantie claim is the most direct empirical support for the RISC-V recommendation, and it has no citation at all. Section 4.1.1 repeats the GPT-3 1280 MWh figure with a vague 'DeepMind' attribution; while a similar number does exist in the literature, the paper does not provide the citation, and the figure is used twice as a core sustainability motivation. Some references in the list are canonical and independently verifiable (Strubell et al., Buolamwini and Gebru, Dwork, Chawla et al.), so this is not a case of wholesale fabrication. But several other references, including [15], [20], and [21], do not correspond to identifiable primary sources, which undermines the survey's reliability as a secondary reference. The concrete test I propose is a straightforward verification of the quantitative claims and the flagged references. Such a check would settle whether the concern lands: if the key performance and energy figures cannot be traced, then the paper's central recommendation is a plausible position with a missing evidentiary base, which is a different epistemic status from a supported survey. Because the reader already classified the paper as UNVERDICTED, my stress-test does not change the verdict; it reinforces it.","tokens_in":13837,"tokens_out":3932,"duration_ms":39892,"concrete_test":"Perform a citation-and-benchmark audit: (1) verify each quantitative claim in §4.1.1 and §4.2.1 (A100 312 TFLOPS/1555 GB/s/400W; GPT-3 1280 MWh; Xuantie 2–3× vs CPU/GPU; TPU 3–5× vs V100; NNP 5×/2×) against vendor datasheets, MLPerf results, or peer-reviewed papers; (2) check the existence and correct bibliographic data of references [15], [20], [21], and any others flagged. If the Xuantie 2–3× claim and the GPT-3 1280 MWh attribution cannot be traced to primary sources, the recommendation loses its empirical grounding and the paper should be marked as an unverified overview, not a reference survey.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is a position/survey. Its central claim—that data, data processing, and model training/inference must be addressed jointly, with RISC-V accelerators as a key enabler—is normative and broad, but its force depends on the factual accuracy of its quantitative exhibits. The weakest link is §4.2.1: the claim that Alibaba Xuantie RISC-V cores achieve 2–3× the performance of CPUs and GPUs in image/video processing is given with no citation, and the SiFive example is similarly source-free. This is not a minor omission: it is the one concrete quantitative justification for the paper's recommendation that RISC-V 'effectively alleviates' AI hardware bottlenecks. The motivation in §4.1.1 is also under-supported: the GPT-3 training figure of 1280 MWh is attributed only to 'research by DeepMind' with no reference, although a verifiable estimate exists in Patterson et al. 2021; correct attribution matters because the number is used repeatedly. Several listed references also fail basic traceability checks (e.g., [20] in a 'Journal of AI and Data Mining' volume/page pattern that does not correspond to a real venue; [21] lists Stoica/Zaharia/Ghodsi in IEEE TCC 2014; [15] has no identifiable primary source). The presence of such references means the survey cannot be used as a reliable secondary source without independent verification. The central claim may still be true, but the current evidence does not establish it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey/position statement on sustainable AI from data and system perspectives. It argues that reducing the environmental impact of AI requires joint attention to three stages — data acquisition, data processing, and AI model training/inference — and it proposes RISC-V-based accelerators, domain-specific architectures, and hardware–software co-optimization as key technical enablers. Each of the three main sections surveys current issues, example solutions, and future challenges. The paper provides no new algorithms, experiments, or derivations; its contribution is a structured narrative and a set of recommendations.","tokens_in":14284,"tokens_out":7140,"duration_ms":61039,"significance":"The paper offers a clear and broad organizational framework that could serve as a useful introduction to sustainability issues in AI, particularly for readers interested in the data-centric perspective. The sections on privacy-preserving data use and data quality are relevant, and the paper explicitly identifies future challenges. However, the paper's scientific value as a secondary source is currently undercut by unsupported quantitative exhibits and unreliable references. Its central recommendation — that RISC-V can 'effectively alleviate' AI performance bottlenecks — rests on uncited performance claims, and several references appear to be misattributed or fabricated. Because the paper is a survey, citation accuracy and factual correctness are load-bearing. The paper does not provide machine-checked proofs, code, or reproducible artifacts; after the factual claims and references are corrected, it could be a serviceable overview.","major_comments":[{"comment":"The paper's central recommendation that RISC-V-based accelerators can effectively alleviate performance bottlenecks rests on the uncited claim that Alibaba's Xuantie RISC-V cores achieve 2–3 times the performance of CPUs and GPUs in image and video processing, and on an equally uncited SiFive example. No primary source is given for either, and the claim is repeated in Section 4.2.3. The authors must provide verifiable references with specifying workload, baseline hardware, precision, and power consumption, or else weaken the claim accordingly.","section":"4.2.1, 4.2.3"},{"comment":"The GPT-3 training energy figure of approximately 1280 MWh is attributed to 'research by DeepMind' with no citation. The published estimate in the literature is from Patterson et al. (2021), which is absent from the reference list. The same figure is repeated in Section 4.3.2. Additionally, the comparison to 'a typical household over ten years' is not correct: 1280 MWh is roughly an order of magnitude larger than a decade of typical U.S. household consumption. Please correct the attribution and the equivalence claim.","section":"4.1.1, 4.3.2"},{"comment":"Several references cannot be traced to published work as cited: [4] is not a known ICDM 2017 paper by Sun and Leskovec; [10] has no identifiable article in IEEE TKDE 2020; [15] is a generic description with no author; [20] cites a journal volume/page that does not match a real article; and [21] credits Stoica, Zaharia, and Ghodsi with a survey on sustainable AI in IEEE TCC 2014, which is not their work. In addition, [8] (IoT survey) is cited for Google's data-center cooling, and [13] cites a 2008 ICALP volume for Dwork's differential privacy, which appeared at ICALP 2006. Because this is a survey paper, the accuracy of the reference list is load-bearing; every citation must be checked against the original source, corrected, or removed.","section":"References [4], [8], [10], [13], [15], [20], [21]"}],"minor_comments":[{"comment":"There are typos and incomplete headings: 'large langrage models' (Abstract), 'F uture Challenges' (Sections 2.3, 3.3, 4.3), 'T raining' in Section 4.1 title, and 'inte lligence' in the running head. These should be cleaned up.","section":"Throughout"},{"comment":"The heading '(By Wentao)' reveals an author name and should be removed; this is not appropriate for a submitted manuscript.","section":"2.2.1 heading"},{"comment":"The statement that the NVIDIA A100 GPU has peak floating-point performance of 312 TFLOPS does not specify precision; the figure corresponds to sparse TF32, not dense FP32, and should be stated with precision.","section":"4.1.1"},{"comment":"The invocation of Fitts's law to explain annotation noise in crowdsourcing is incorrect: Fitts's law is a model of human pointing movement, not the speed-accuracy trade-off in labeling. This should be corrected or removed.","section":"2.2.1"},{"comment":"Several 'example solutions' (e.g., Google's data-center cooling, Waymo/NVIDIA synthetic data, Google's active learning pipelines) are described without specific citations; add primary references or mark them as illustrative.","section":"2.2.2, 2.2.3, 3.2.4"},{"comment":"The references are in an inconsistent format (e.g., [11] and [15] are not scholarly citations) and should be harmonized.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript needs a full verification of its reference list before it can be considered for publication. I would ask the editor to require the authors to submit DOIs or URLs for all citations, and to treat the multiple untraceable entries as an integrity issue. The '(By Wentao)' header also suggests the paper was not prepared for anonymous review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one for what it is: a competent but shallow survey, and not a reliable secondary source in its current form. The three-way split—data acquisition, data processing, training/inference—is a sensible frame, and the paper does a fair job of listing familiar techniques (active learning, SMOTE, federated learning, AutoAugment) with worked example solutions. That is the extent of the novelty; there is no new method, dataset, or analysis.\n\nThe soft spots are real and load-bearing. The main unsupported numbers are the GPT-3 training energy figure (1280 MWh, attributed to \"research by DeepMind\" with no citation) and the claim in §4.2.1 that Alibaba's Xuantie RISC-V cores achieve 2–3× the performance of CPUs and GPUs on image/video processing, also with no citation. Those numbers are the principal evidence for the paper's RISC-V recommendation, so the recommendation rests on sand. A verifiable GPT-3 estimate exists (Patterson et al., 2021), but it is not cited.\n\nWorse, several references fail basic traceability checks: [20] (Journal of AI and Data Mining), [21] (Stoica/Zaharia/Ghodsi in IEEE TCC 2014), [15] (HealthDataSpace), and [4] (Sun/Leskovec on mining non-textual data) do not resolve to plausible primary sources. That means the survey cannot be used as a map to the literature without independent verification, which defeats the purpose of a survey.\n\nThe central normative claim—that sustainability needs joint attention to data and systems—is plausible and widely accepted, so the paper's direction is not wrong. But the evidence presented does not establish it, and the citation integrity issues are disqualifying for a review article.\n\nWho is this for? Someone who wants a quick, high-level orientation to the buzzwords in sustainable AI and does not mind double-checking every source. It is not for a researcher trying to cite a trustworthy overview.\n\nIf this crossed my desk, I would desk-reject and invite resubmission only after the authors fix the references and substantiate the quantitative claims. It does not deserve referee time in its current state.","headline":"A readable but shallow survey whose unsupported numbers and shaky reference list make it unreliable as a secondary source in its current form.","tokens_in":14676,"tokens_out":2978,"would_cite":false,"duration_ms":27604,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that sustainable AI requires coordinated improvements in data acquisition, data processing, and model training, with open hardware like RISC-V as a key lever.","keywords":["sustainable AI","data acquisition","data processing","energy-efficient computing","RISC-V","domain-specific architecture","hardware-software co-optimization","data-centric AI"],"falsifier":"Run a controlled benchmark comparing a current RISC-V AI accelerator against a mainstream CPU and GPU on a representative set of deep-learning workloads (e.g., ResNet-50 inference and GPT-class language-model inference), measuring both throughput and energy per inference; if the RISC-V system does not achieve better energy efficiency than the CPU or GPU, the paper's central hardware recommendation is undermined.","tokens_in":13675,"feed_emoji":"🌱","tokens_out":2627,"duration_ms":25754,"temperature":0.7,"pith_summary":"The paper makes the case that reducing the environmental footprint of AI is not a single optimization problem but a systems problem spanning data acquisition, data processing, and model training and inference. It surveys the current issues in each area—energy-hungry data collection, noisy and imbalanced datasets, and performance bottlenecks in existing hardware—and argues that targeted techniques in each can help. The strongest claim is that open instruction-set architectures like RISC-V, combined with domain-specific designs and hardware-software co-optimization, can meaningfully alleviate the performance and energy bottlenecks of AI workloads. A sympathetic reader would take away that sustainability should be a design criterion across the whole AI stack, not an afterthought at the model level.","feed_headline":"AI can go greener by fixing data and hardware together","feed_subtitle":"A review ties AI's environmental toll to data handling and open, customizable chips like RISC-V.","key_machinery":"The paper's central organizing device is the three-phase lifecycle of AI systems—data acquisition, data processing, and model training/inference—treated as a whole. The load-bearing technical mechanisms are: RISC-V as an open instruction-set architecture that permits customized accelerators; domain-specific architectures that specialize hardware for particular neural-network operations; and hardware-software co-optimization through compilers like TensorFlow XLA that tune computation graphs to the underlying hardware. On the data side, the key mechanisms are active learning, synthetic data generation, automated cleaning, and privacy-preserving techniques, all of which cut wasted computation or data collection.","core_discovery":"The paper asserts that sustainable AI can be achieved by jointly addressing three pillars: data acquisition, data processing, and AI model training and inference. For data, it argues that cost-effective collection, active learning, synthetic data, and privacy-preserving techniques such as federated learning and differential privacy reduce both energy and waste. For processing, automated cleaning, feature engineering, and intelligent augmentation improve efficiency. For the model side, it claims that RISC-V-based AI accelerators, domain-specific architectures, and hardware-software co-optimization can break through current performance bottlenecks, with the paper stating that RISC-V customization and co-optimization 'can be effectively alleviated' the performance bottlenecks in AI hardware architectures.","pith_inferences":["The paper implicitly suggests that embodied carbon from manufacturing new accelerators could offset operational savings, an issue it does not quantify; a full life-cycle assessment of RISC-V-based systems would be a natural follow-up.","It is reasonable to expect that combining the data-centric and hardware-centric recommendations would yield multiplicative rather than additive energy savings, since smaller, cleaner datasets require less compute and thus less hardware capacity.","The claimed performance advantages of RISC-V are based on vendor or anecdotal benchmarks; a public, standardized benchmark suite comparing RISC-V accelerators against CPUs and GPUs on realistic AI workloads would test whether the recommendation transfers beyond the paper's examples.","The paper's emphasis on non-textual data like acoustic and sensor data points toward an underexplored opportunity: efficient multimodal processing could enable AI applications that monitor environmental sustainability directly, creating a positive feedback loop."],"forward_implications":["Adopting the paper's recommended data-centric techniques would reduce the amount of data that must be collected and labeled, lowering the energy spent on acquisition and preprocessing.","RISC-V-based accelerators, if the cited performance figures hold, could deliver 2–3 times the performance of CPUs and GPUs for image and video processing at similar power, making edge inference more sustainable.","Hardware-software co-optimization, such as compiler-level tuning for custom accelerators, would allow AI frameworks to squeeze more useful computation per watt, especially for large-scale training.","The framework implies that sustainability metrics should be attached to data pipelines and hardware choices, not just to model flops, to guide future design decisions.","Future AI hardware development should prioritize flexibility and customizability—exemplified by RISC-V—to adapt to diverse AI workloads without sacrificing energy efficiency."],"supporting_citations":[{"why":"Quantifies the energy and carbon cost of training large NLP models, establishing the motivation for sustainable AI.","marker":"[1]"},{"why":"Introduces federated learning as a privacy-preserving technique to use private data without centralizing it.","marker":"[12]"},{"why":"Defines differential privacy, a key mechanism the paper relies on for privacy-preserving data utilization.","marker":"[13]"},{"why":"Provides SMOTE, the synthetic minority over-sampling technique central to the paper's imbalanced-data solutions.","marker":"[31]"},{"why":"Describes AutoAugment, an automated augmentation method cited as an intelligent way to generate realistic training data.","marker":"[34]"},{"why":"Surveys data cleaning challenges and emerging approaches, grounding the paper's claims about automated cleaning.","marker":"[22]"},{"why":"Discusses truth inference for crowdsourced annotations, supporting the paper's cost-effective data acquisition proposals.","marker":"[6]"},{"why":"References energy-efficient data-center optimization, an example the paper uses for improving AI system energy efficiency.","marker":"[8]"}],"fun_headline_variants":["Greener AI needs better data and open chips like RISC-V","RISC-V chips and smart data could cut AI's carbon footprint","Data efficiency plus RISC-V hardware: the path to sustainable AI","Sustainable AI: fix data pipelines and embrace RISC-V chips"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's case for RISC-V as a key enabler rests on performance and energy figures—such as the claimed 2–3 times speedup of Alibaba's Xuantie cores over CPUs and GPUs—being accurate and representative when the paper cites no primary source for them.","fun_headline_variants_meta":{"raw":{"variants":["Greener AI needs better data and open chips like RISC-V","RISC-V chips and smart data could cut AI's carbon footprint","Data efficiency plus RISC-V hardware: the path to sustainable AI","Sustainable AI: fix data pipelines and embrace RISC-V chips"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000385,"raw_usage":{"total_tokens":1936,"prompt_tokens":746,"completion_tokens":1190,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":362,"completion_tokens_details":{"reasoning_tokens":1116}},"tokens_in":362,"tokens_out":1190,"duration_ms":8875,"temperature":1.0,"reasoning_tokens":1116,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:39:41.135685+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled benchmark comparing a current RISC-V AI accelerator against a mainstream CPU and GPU on a representative set of deep-learning workloads (e.g., ResNet-50 inference and GPT-class language-model inference), measuring both throughput and energy per inference; if the RISC-V system does not achieve better energy efficiency than the CPU or GPU, the paper's central hardware recommendation is undermined.","supporting_citations":[{"cited_title":"Energy and policy considerations for deep learning in nlp","cited_arxiv_id":null,"evidence_quote":"Quantifies the energy and carbon cost of training large NLP models, establishing the motivation for sustainable AI."},{"cited_title":"Communication-eﬃcient learning of deep net- works from decentralized data","cited_arxiv_id":null,"evidence_quote":"Introduces federated learning as a privacy-preserving technique to use private data without centralizing it."},{"cited_title":"Diﬀerential privacy","cited_arxiv_id":null,"evidence_quote":"Defines differential privacy, a key mechanism the paper relies on for privacy-preserving data utilization."},{"cited_title":"Smote: synthetic minority over-sampling technique","cited_arxiv_id":null,"evidence_quote":"Provides SMOTE, the synthetic minority over-sampling technique central to the paper's imbalanced-data solutions."},{"cited_title":"Autoaugment: Learning augmenta- tion strategies from data","cited_arxiv_id":null,"evidence_quote":"Describes AutoAugment, an automated augmentation method cited as an intelligent way to generate realistic training data."},{"cited_title":"Data clean- ing: Overview and emerging challenges","cited_arxiv_id":null,"evidence_quote":"Surveys data cleaning challenges and emerging approaches, grounding the paper's claims about automated cleaning."},{"cited_title":"Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment , 2017, 10(5):541–552","cited_arxiv_id":null,"evidence_quote":"Discusses truth inference for crowdsourced annotations, supporting the paper's cost-effective data acquisition proposals."},{"cited_title":"In- ternet of things (iot): A vision, architectural ele- ments, and future directions","cited_arxiv_id":null,"evidence_quote":"References energy-efficient data-center optimization, an example the paper uses for improving AI system energy efficiency."}],"review_version":1}