Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Data and System Perspectives of Sustainable Artificial Intelligence

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper argues that sustainable AI requires coordinated improvements in data acquisition, data processing, and model training, with open hardware like RISC-V as a key lever.

desk verdict A readable but shallow survey whose unsupported numbers and shaky reference list make it unreliable as a secondary source in its current form. read the letter →

arxiv 2501.07487 v1 pith:JNOMROLO submitted 2025-01-13 cs.AI

classification cs.AI
keywords sustainableAIdataacquisitionprocessingenergy-efficientcomputingRISC-Vdomain-specificarchitecturehardware-softwareco-optimizationdata-centric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper makes the case that reducing the environmental footprint of AI is not a single optimization problem but a systems problem spanning data acquisition, data processing, and model training and inference. It surveys the current issues in each area—energy-hungry data collection, noisy and imbalanced datasets, and performance bottlenecks in existing hardware—and argues that targeted techniques in each can help. The strongest claim is that open instruction-set architectures like RISC-V, combined with domain-specific designs and hardware-software co-optimization, can meaningfully alleviate the performance and energy bottlenecks of AI workloads. A sympathetic reader would take away that sustainability should be a design criterion across the whole AI stack, not an afterthought at the model level.

What carries the argument

The paper's central organizing device is the three-phase lifecycle of AI systems—data acquisition, data processing, and model training/inference—treated as a whole. The load-bearing technical mechanisms are: RISC-V as an open instruction-set architecture that permits customized accelerators; domain-specific architectures that specialize hardware for particular neural-network operations; and hardware-software co-optimization through compilers like TensorFlow XLA that tune computation graphs to the underlying hardware. On the data side, the key mechanisms are active learning, synthetic data generation, automated cleaning, and privacy-preserving techniques, all of which cut wasted computation or data collection.

What would settle it

Run a controlled benchmark comparing a current RISC-V AI accelerator against a mainstream CPU and GPU on a representative set of deep-learning workloads (e.g., ResNet-50 inference and GPT-class language-model inference), measuring both throughput and energy per inference; if the RISC-V system does not achieve better energy efficiency than the CPU or GPU, the paper's central hardware recommendation is undermined.

Watch

Extended reading notes

Core claim

The paper asserts that sustainable AI can be achieved by jointly addressing three pillars: data acquisition, data processing, and AI model training and inference. For data, it argues that cost-effective collection, active learning, synthetic data, and privacy-preserving techniques such as federated learning and differential privacy reduce both energy and waste. For processing, automated cleaning, feature engineering, and intelligent augmentation improve efficiency. For the model side, it claims that RISC-V-based AI accelerators, domain-specific architectures, and hardware-software co-optimization can break through current performance bottlenecks, with the paper stating that RISC-V customization and co-optimization 'can be effectively alleviated' the performance bottlenecks in AI hardware architectures.

Load-bearing premise

The paper's case for RISC-V as a key enabler rests on performance and energy figures—such as the claimed 2–3 times speedup of Alibaba's Xuantie cores over CPUs and GPUs—being accurate and representative when the paper cites no primary source for them.

Editorial extensions

If this is right

  • Adopting the paper's recommended data-centric techniques would reduce the amount of data that must be collected and labeled, lowering the energy spent on acquisition and preprocessing.
  • RISC-V-based accelerators, if the cited performance figures hold, could deliver 2–3 times the performance of CPUs and GPUs for image and video processing at similar power, making edge inference more sustainable.
  • Hardware-software co-optimization, such as compiler-level tuning for custom accelerators, would allow AI frameworks to squeeze more useful computation per watt, especially for large-scale training.
  • The framework implies that sustainability metrics should be attached to data pipelines and hardware choices, not just to model flops, to guide future design decisions.
  • Future AI hardware development should prioritize flexibility and customizability—exemplified by RISC-V—to adapt to diverse AI workloads without sacrificing energy efficiency.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper implicitly suggests that embodied carbon from manufacturing new accelerators could offset operational savings, an issue it does not quantify; a full life-cycle assessment of RISC-V-based systems would be a natural follow-up.
  • It is reasonable to expect that combining the data-centric and hardware-centric recommendations would yield multiplicative rather than additive energy savings, since smaller, cleaner datasets require less compute and thus less hardware capacity.
  • The claimed performance advantages of RISC-V are based on vendor or anecdotal benchmarks; a public, standardized benchmark suite comparing RISC-V accelerators against CPUs and GPUs on realistic AI workloads would test whether the recommendation transfers beyond the paper's examples.
  • The paper's emphasis on non-textual data like acoustic and sensor data points toward an underexplored opportunity: efficient multimodal processing could enable AI applications that monitor environmental sustainability directly, creating a positive feedback loop.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is a survey/position statement on sustainable AI from data and system perspectives. It argues that reducing the environmental impact of AI requires joint attention to three stages — data acquisition, data processing, and AI model training/inference — and it proposes RISC-V-based accelerators, domain-specific architectures, and hardware–software co-optimization as key technical enablers. Each of the three main sections surveys current issues, example solutions, and future challenges. The paper provides no new algorithms, experiments, or derivations; its contribution is a structured narrative and a set of recommendations.

Significance. The paper offers a clear and broad organizational framework that could serve as a useful introduction to sustainability issues in AI, particularly for readers interested in the data-centric perspective. The sections on privacy-preserving data use and data quality are relevant, and the paper explicitly identifies future challenges. However, the paper's scientific value as a secondary source is currently undercut by unsupported quantitative exhibits and unreliable references. Its central recommendation — that RISC-V can 'effectively alleviate' AI performance bottlenecks — rests on uncited performance claims, and several references appear to be misattributed or fabricated. Because the paper is a survey, citation accuracy and factual correctness are load-bearing. The paper does not provide machine-checked proofs, code, or reproducible artifacts; after the factual claims and references are corrected, it could be a serviceable overview.

major comments (3)
  1. [4.2.1, 4.2.3] The paper's central recommendation that RISC-V-based accelerators can effectively alleviate performance bottlenecks rests on the uncited claim that Alibaba's Xuantie RISC-V cores achieve 2–3 times the performance of CPUs and GPUs in image and video processing, and on an equally uncited SiFive example. No primary source is given for either, and the claim is repeated in Section 4.2.3. The authors must provide verifiable references with specifying workload, baseline hardware, precision, and power consumption, or else weaken the claim accordingly.
  2. [4.1.1, 4.3.2] The GPT-3 training energy figure of approximately 1280 MWh is attributed to 'research by DeepMind' with no citation. The published estimate in the literature is from Patterson et al. (2021), which is absent from the reference list. The same figure is repeated in Section 4.3.2. Additionally, the comparison to 'a typical household over ten years' is not correct: 1280 MWh is roughly an order of magnitude larger than a decade of typical U.S. household consumption. Please correct the attribution and the equivalence claim.
  3. [References [4], [8], [10], [13], [15], [20], [21]] Several references cannot be traced to published work as cited: [4] is not a known ICDM 2017 paper by Sun and Leskovec; [10] has no identifiable article in IEEE TKDE 2020; [15] is a generic description with no author; [20] cites a journal volume/page that does not match a real article; and [21] credits Stoica, Zaharia, and Ghodsi with a survey on sustainable AI in IEEE TCC 2014, which is not their work. In addition, [8] (IoT survey) is cited for Google's data-center cooling, and [13] cites a 2008 ICALP volume for Dwork's differential privacy, which appeared at ICALP 2006. Because this is a survey paper, the accuracy of the reference list is load-bearing; every citation must be checked against the original source, corrected, or removed.
minor comments (6)
  1. [Throughout] There are typos and incomplete headings: 'large langrage models' (Abstract), 'F uture Challenges' (Sections 2.3, 3.3, 4.3), 'T raining' in Section 4.1 title, and 'inte lligence' in the running head. These should be cleaned up.
  2. [2.2.1 heading] The heading '(By Wentao)' reveals an author name and should be removed; this is not appropriate for a submitted manuscript.
  3. [4.1.1] The statement that the NVIDIA A100 GPU has peak floating-point performance of 312 TFLOPS does not specify precision; the figure corresponds to sparse TF32, not dense FP32, and should be stated with precision.
  4. [2.2.1] The invocation of Fitts's law to explain annotation noise in crowdsourcing is incorrect: Fitts's law is a model of human pointing movement, not the speed-accuracy trade-off in labeling. This should be corrected or removed.
  5. [2.2.2, 2.2.3, 3.2.4] Several 'example solutions' (e.g., Google's data-center cooling, Waymo/NVIDIA synthetic data, Google's active learning pipelines) are described without specific citations; add primary references or mark them as illustrative.
  6. [References] The references are in an inconsistent format (e.g., [11] and [15] are not scholarly citations) and should be harmonized.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey with no derivation, no fitted parameters, and no load-bearing self-citations.

full rationale

This paper is a narrative literature survey and position statement rather than a derivation. It does not fit parameters to data, does not define its concepts in terms of its conclusions, and does not present original predictive results that could be forced by construction. The central recommendations, such as jointly addressing data acquisition, data processing, and model training/inference while leveraging RISC-V-based accelerators, are supported by external citations and example solutions, not by the authors' own prior results. The section header '(By Wentao)' in Section 2.2.1 is an authorship annotation, not a circular argument. The quantitative claims that could be load-bearing for the RISC-V recommendation, including the 1280 MWh GPT-3 figure in Section 4.1.1 and the 2-3x Xuantie performance claim in Section 4.2.1, lack primary citations or verifiable sourcing, but that is a correctness and evidence-quality concern, not circularity. No equation is reused as a prediction, no fitted parameter is renamed as a result, and no uniqueness or foundational theorem is imported from the authors' own prior work. Accordingly, the circularity burden is minimal and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces a conceptual taxonomy but no new physical or formal entities, so no invented entities are needed. It also provides no fitted parameters or quantitative model, so there are no free parameters.

assumptions (2)
  • domain assumption The cited works are real and accurately support the claims in the text.
    The review's assertions about energy use, performance gains, and data techniques rely on the 34 references, several of which are unverifiable or likely fabricated, so the reliability of the entire survey is at risk.
  • domain assumption The described technologies (active learning, SMOTE, federated learning, differential privacy, RISC-V accelerators) function as portrayed in the cited literature.
    The paper presents these as effective solutions without original evaluation, assuming the cited sources are trustworthy and representative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data and System Perspectives of Sustainable Artificial Intelligence." pith.science (2026). https://pith.science/paper/JNOMROLO

@misc{pith2026250107487,
  author       = {Pith},
  title        = {Pith review of: Data and System Perspectives of Sustainable Artificial Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JNOMROLO}},
  note         = {Machine review of arXiv:2501.07487}
}
read the original abstract

Sustainable AI is a subfield of AI for concerning developing and using AI systems in ways of aiming to reduce environmental impact and achieve sustainability. Sustainable AI is increasingly important given that training of and inference with AI models such as large langrage models are consuming a large amount of computing power. In this article, we discuss current issues, opportunities and example solutions for addressing these issues, and future challenges to tackle, from the data and system perspectives, related to data acquisition, data processing, and AI model training and inference.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

    cs.CL 2025-06 conditional novelty 6.0 of 10

    StoryBench introduces a branching interactive-fiction benchmark with immediate-feedback and self-recovery modes, and shows that current LLMs fail at long-term memory tasks, especially self-correction.

Reference graph

Works this paper leans on

34 extracted references · 32 canonical work pages · cited by 1 Pith paper

  1. [20]

    The standardization of non-textual data for ai systems: Challenges and opportunities

    Herzog T, Ochoa M T L M. The standardization of non-textual data for ai systems: Challenges and opportunities. Journal of AI and Data Mining , 2021, 32(5):438–450

  2. [21]

    A survey of sustain- able ai scalability techniques

    Stoica I, Zaharia M, Ghodsi A. A survey of sustain- able ai scalability techniques. IEEE Transactions on Cloud Computing , 2014, 2(3):231–243

  3. [15]

    The healthdataspace project: En- abling secure cross-border health data sharing for ai applications

    Consortium H. The healthdataspace project: En- abling secure cross-border health data sharing for ai applications. European Commission, 2021

  4. [4]

    Mining non-textual data for deep learning models

    Sun X, Leskovec J. Mining non-textual data for deep learning models. Proceedings of the IEEE In- ternational Conference on Data Mining (ICDM) , 2017, pp. 473–482

  5. [10]

    Synthetic data genera- tion for privacy-preserving machine learning

    Franklin A, Cook J D. Synthetic data genera- tion for privacy-preserving machine learning. IEEE Transactions on Knowledge and Data Engineering , 2020, 32(12):2347–2359

  6. [8]

    In- ternet of things (iot): A vision, architectural ele- ments, and future directions

    Gubbi J, Buyya R, Marusic S, Palaniswami M. In- ternet of things (iot): A vision, architectural ele- ments, and future directions. Future Generation Computer Systems , 2013, 29(7):1645–1660

  7. [13]

    Differential privacy

    Dwork C. Differential privacy. Proceedings of the 33rd International Colloquium on Automata, Lan- guages and Programming , 2008, pp. 1–12

  8. [1]

    Energy and policy considerations for deep learning in nlp

    Strubell R, Ganesh A, McCallum A. Energy and policy considerations for deep learning in nlp. Pro- ceedings of the 57th Annual Meeting of the Asso- 16 J. Comput. Sci. & Technol. ciation for Computational Linguistics , 2019, pp. 3645–3650

Show all 34 references
  1. [2]

    Gender shades: Intersec- tional accuracy disparities in commercial gender classification

    Buolamwini J, Gebru T. Gender shades: Intersec- tional accuracy disparities in commercial gender classification. Proceedings of the 1st Conference on Fairness, Accountability and Transparency , 2018, pp. 77–91

  2. [3]

    Mem- bership inference attacks against machine learning models

    Shokri R, Stronati M, Song C, Shmatikov V. Mem- bership inference attacks against machine learning models. 2017 IEEE Symposium on Security and Privacy (SP) , 2017, pp. 3–18

  3. [5]

    Robust de- anonymization of large sparse datasets

    Narayanan A, Shmatikov V. Robust de- anonymization of large sparse datasets. In Pro- ceedings of the 2008 IEEE Symposium on Security and Privacy , SP ’08, 2008, pp. 111–125

  4. [6]

    Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment , 2017, 10(5):541–552

    Zheng Y, Li G, Li Y, Shan C, Cheng R. Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment , 2017, 10(5):541–552

  5. [7]

    An iterative and re- weighting framework for rejection and uncertainty resolution in crowdsourcing

    Xie S, Fan W, Yu P S. An iterative and re- weighting framework for rejection and uncertainty resolution in crowdsourcing. In Proceedings of the 2012 SIAM International Conference on Data Mining, 2012, pp. 1107–1118

  6. [9]

    Ethereum: A Secure Decentralized Gen- eralized Transaction Ledger

    Wood G. Ethereum: A Secure Decentralized Gen- eralized Transaction Ledger . Ethereum Project Yellow Paper, 2014

  7. [11]

    Waymo’s approach to self-driving car de- velopment

    Waymo. Waymo’s approach to self-driving car de- velopment. Waymo Blog , 2019

  8. [12]

    Communication-efficient learning of deep net- works from decentralized data

    McMahan B, Moore E, Ramage D. Communication-efficient learning of deep net- works from decentralized data. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS) , 2017, 54:1273–1282

  9. [14]

    The international data spaces initiative: A blueprint for secure data sharing

    Bohnenberger M, Behrens L D P, Flemming R S. The international data spaces initiative: A blueprint for secure data sharing. Interna- tional Journal of Information Management , 2021, 57:102325

  10. [16]

    Ai for environmental monitor- ing: Applications and challenges

    Xu J, Liu Q, Liu Y. Ai for environmental monitor- ing: Applications and challenges. Environmental Science & Technology, 2020, 54(9):5594–5603

  11. [17]

    Efficient fully homomorphic encryption from (standard) lwe

    Brakerski Z, Halevi S, Smart N P. Efficient fully homomorphic encryption from (standard) lwe. Proceedings of the 35th Annual International Conference on the Theory and Applications of Data & System Perspectives of Sustainable AI 17 Cryptographic Techniques (EUROCRYPT) , 2014, pp. 1–17

  12. [18]

    Homomorphic encryption for privacy-preserving machine learning: A survey

    Chen X, Liu S, Xiong H. Homomorphic encryption for privacy-preserving machine learning: A survey. ACM Computing Surveys , 2020, 53(3):1–37

  13. [19]

    Privacy, surveillance, and public trust

    Regan P M. Privacy, surveillance, and public trust. Internet Policy Review , 2015, 4(4)

  14. [22]

    Data clean- ing: Overview and emerging challenges

    Chu X, Ilyas I F, Krishnan S, Wang J. Data clean- ing: Overview and emerging challenges. In Pro- ceedings of the 2016 international conference on management of data , 2016, pp. 2201–2206

  15. [23]

    Data management in machine learning: Challenges, techniques, and systems

    Kumar A, Boehm M, Yang J. Data management in machine learning: Challenges, techniques, and systems. In Proceedings of the 2017 ACM Interna- tional Conference on Management of Data , 2017, pp. 1717–1722

  16. [24]

    Learning from imbalanced data

    He H, Garcia E A. Learning from imbalanced data. IEEE Transactions on knowledge and data engi- neering, 2009, 21(9):1263–1284

  17. [25]

    In defense of core-set: A density- aware core-set selection for active learning

    Kim Y, Shin B. In defense of core-set: A density- aware core-set selection for active learning. In Pro- ceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 804–812

  18. [26]

    A survey of data aug- mentation approaches for nlp

    Feng S Y, Gangal V, Wei J, Chandar S, Vosoughi S, Mitamura T, Hovy E. A survey of data aug- mentation approaches for nlp. In Findings of the Association for Computational Linguistics: ACL- IJCNLP 2021 , 2021, pp. 968–988

  19. [27]

    Alphaclean: Automatic gen- eration of data cleaning pipelines

    Krishnan S, Wu E. Alphaclean: Automatic gen- eration of data cleaning pipelines. arXiv preprint arXiv:1904.11827, 2019

  20. [28]

    Log-based anomaly detection with deep learning: How far are we? In Proceed- ings of the 44th international conference on soft- ware engineering, 2022, pp

    Le V H, Zhang H. Log-based anomaly detection with deep learning: How far are we? In Proceed- ings of the 44th international conference on soft- ware engineering, 2022, pp. 1356–1367

  21. [29]

    Using openrefine

    Verborgh R, De Wilde M. Using openrefine. Packt Publishing Ltd, 2013

  22. [30]

    Deep feature syn- thesis: Towards automating data science endeav- ors

    Kanter J M, Veeramachaneni K. Deep feature syn- thesis: Towards automating data science endeav- ors. In 2015 IEEE international conference on data science and advanced analytics (DSAA) , 2015, pp. 1–10

  23. [31]

    Smote: synthetic minority over-sampling technique

    Chawla N V, Bowyer K W, Hall L O, Kegelmeyer W P. Smote: synthetic minority over-sampling technique. Journal of artificial intelligence re- search, 2002, 16:321–357

  24. [32]

    Training cost-sensitive neural networks with methods addressing the class imbal- ance problem

    Zhou Z H, Liu X Y. Training cost-sensitive neural networks with methods addressing the class imbal- ance problem. IEEE Transactions on knowledge and data engineering , 2005, 18(1):63–77

  25. [33]

    Active learning for convolu- tional neural networks: A core-set approach

    Sener O, Savarese S. Active learning for convolu- tional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489, 2017

  26. [34]

    Autoaugment: Learning augmenta- tion strategies from data

    Cubuk E D, Zoph B, Mane D, Vasudevan V, Le Q V. Autoaugment: Learning augmenta- tion strategies from data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 113–123

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.