{"id":"1c0f4f00-e639-410a-af65-ed6f8cbc69d5","arxiv_id":"2502.09651","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors present AI-VERDE, a university LLM gateway, and report a six-month pilot with 78 users, 5 courses, about 110 million tokens, and 97,658 API calls.","lead":"AI-VERDE is a campus-wide platform at the University of Arizona that gives students, faculty, and staff one login and API key to access many different AI language models, both free open-source and paid commercial ones. The paper describes the platform's design and reports early usage from a six-month pilot at a large public university.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'egalitarian and lower-cost' claim is not backed by the cost comparison the paper promises; no budget, hardware, or per-token cost figures are reported, so the paper cannot establish that AI-VERDE is cheaper or more equitable than commercial APIs.","rationale":"The paper is a systems-and-deployment paper, not a modeling paper. Its components (vLLM, LiteLLM, Weaviate, CILogon) are standard open-source pieces, so the claimed novelty and value rest on the integration as an egalitarian, low-cost gateway for teaching and research. Contribution 3 explicitly promises a comparative cost study, and the abstract highlights enabling activities that would 'typically require substantial budgets' for commercial services. No such comparison is delivered: there are no dollar figures, no per-token rates, no hardware amortization, and no baseline against the commercial APIs the paper names. This is the single most load-bearing gap because it is the part of the central claim that is both quantitative and absent. The pilot usage numbers are useful, but they do not by themselves show cost-effectiveness; indeed, the dominance of self-hosted tokens (109.76M of 110.68M) means the economic question is whether fixed hosting costs beat per-token commercial pricing, which cannot be answered from Table 2. The reader's concern about generalizing from 78 pilot users is related but secondary: even a very large pilot would not establish 'egalitarian' unless the economics are favorable. I would keep the paper conditional rather than reject it, because the deployment is real, the usage statistics are concrete, and the missing cost analysis is straightforwardly fixable if the data exist.","tokens_in":14111,"tokens_out":3714,"duration_ms":36288,"concrete_test":"Ask the authors to release the pilot's actual cost ledger: hardware acquisition and amortization, cloud and proxy spend, staff and consulting time, and electricity, then compute cost per million tokens and per active user for AI-VERDE. Compare that with contemporaneous commercial list prices (e.g., GPT-4o and GPT-4o mini, Claude) for the same 110.68M-token workload (75.65M prompt plus 35.03M completion tokens) and with Anvil's grant-funded cost. If AI-VERDE's all-in cost per token or per user is not below the commercial baseline, the 'egalitarian/lower-cost' claim in Contribution 3 and Section 3.4 should be withdrawn or substantially weakened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Contribution 3 (Section 1) promises 'a comparative study with other commercial options' showing 'a much lower cost egalitarian gateway,' and Section 3.4 claims AI-VERDE ensures 'minimal operational costs' via open-source components and partnerships. Yet no cost comparison appears anywhere. Section 3.6 and Appendix F mention only that hosting LLMs often exceeds $50,000; they give no AI-VERDE hardware cost, no personnel/operating cost, no per-token or per-user cost, and no baseline against OpenAI or Anvil pricing. The six-month pilot used 97,658 API calls and 110.68M tokens (Table 2), but from Table 2 alone one cannot tell whether this workload would have been cheaper on a commercial API. In fact, only 0.919M of the 110.68M tokens were relayed to third-party proxies; almost all load was self-hosted, which shifts cost from per-token fees to fixed GPU/cluster expenses that are never disclosed. Since the egalitarian-access claim rests on affordability, the absence of cost data is the load-bearing gap: a deployment can be feature-complete and still fail the paper's central value proposition if it is not actually cheaper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AI-VERDE, an LLM platform-as-a-service deployed at the University of Arizona that unifies access to self-hosted open models (served via vLLM), commercial model APIs and research services such as AnvilGPT behind a LiteLLM proxy, and adds course/group management, per-course budget controls, native RAG support through a managed Weaviate instance, and a conversational web interface with CILogon authentication. The authors report a six-month pilot with 78 users across 5 courses and 10 research projects, totaling 97,658 API calls and 110.68 million tokens, and they argue from a 372-participant qualitative survey that the platform addresses privacy, intellectual property, cost, and usability barriers that block LLM adoption in higher education. The paper claims that AI-VERDE is the first platform to address both instructional and research LLM needs within a higher education institutional framework and that it provides a lower-cost 'egalitarian gateway' relative to commercial options.","tokens_in":14352,"tokens_out":4223,"duration_ms":37767,"significance":"If the cost and adoption claims were substantiated with data, this systems contribution would be directly useful to university computing centers and comparable institutional service providers. The architecture description is concrete and replicable, the deployment statistics are internally consistent, and Table 2 usefully disaggregates self-hosted versus proxy token volumes and API calls. The authors' explicit acknowledgments of limitations (dedicated hardware constraints, the privacy boundary for proxied commercial traffic, and residual model bias) are commendable and improve the paper's credibility. However, the central value proposition of an egalitarian, lower-cost gateway is asserted rather than demonstrated: no cost comparison is provided, the pilot population is small and self-selected, and strong claims about reduced hallucination and HIPAA-grade privacy are made without measurement. As it stands, the paper is best read as a deployment blueprint and requirements catalog, not as an evaluated solution to the affordability problem it highlights.","major_comments":[{"comment":"The paper promises a comparative study against commercial options showing a 'much lower cost egalitarian gateway' (Section 1, Contribution 3), and Section 3.4 claims 'minimal operational costs' through open-source components and partnerships, yet no cost comparison appears anywhere in the manuscript. Section 3.6 and Appendix F contain only the statement that hosting LLMs often exceeds $50,000; there are no figures for AI-VERDE's hardware, personnel, per-token, or per-user costs, and no baseline against OpenAI, Gemini, or Anvil pricing. Table 2 shows that 109.76M of the 110.68M tokens were self-hosted, which shifts cost from per-token fees to fixed infrastructure expense, but that fixed expense is never disclosed. Because the egalitarian-access claim rests on affordability, this missing cost evidence is load-bearing and must be closed with actual deployment costs or a clearly scoped cost model.","section":"Section 1, Contribution 3; Section 3.4; Section 3.6; Appendix F"},{"comment":"The claims of 'HIPAA-grade' privacy via the Soteria integration and of 'minimal hallucination' via a specially engineered prompt are made without any audit, compliance certification, or evaluation. No HIPAA compliance assessment is cited, and the hallucination discussion provides no test set, baseline model, or quantitative result such that a reader could verify the claim. These claims should either be removed or replaced with measurable evidence, and the paper should state what certification or review the Soteria integration has actually undergone.","section":"Section 3.5 and Appendix B.2"},{"comment":"The generalization that AI-VERDE has 'huge potential for adoption in academic settings' is not supported by the evidence presented. The pilot involved 78 users, a number the authors themselves describe as 'relatively low' for an institution of this scale, and the survey in Appendix A recruited through sign-up flyers, snowball referrals, and an open campus call, so the 372 participants cannot be treated as representative of the university community. The paper should either reframe the conclusions as pilot-scale observations or provide a sampling plan, response rates, and demographic coverage information that would support population-level claims about egalitarian access.","section":"Section 4.2 and Appendix A"}],"minor_comments":[{"comment":"The text says 'more than 110 tokens were passed' but the table to which it refers reports 110.68 million tokens; the 'million' is missing.","section":"Section 4.2"},{"comment":"The phrase 'an higher education institutional framework' should be 'a higher education institutional framework'; similar article errors appear elsewhere (e.g., 'an higher education institution' in the abstract).","section":"Abstract and Section 1"},{"comment":"Items 7 through 9 are all labeled 'Harwood Lab' but one of the URLs points to Joseph Bonito's page, so the labels do not match the cited URLs; the list also contains a duplicated label for Harwood across items 7 and 8.","section":"Appendix C"},{"comment":"Appendix F is titled 'Cost details' but contains only the same single sentence as Section 3.6; either expand it with the actual cost analysis promised in the introduction or rename it to avoid implying that cost data are provided.","section":"Appendix F"},{"comment":"The text cites Touvron et al. (2023) for Llama 3.2, but that reference is the Llama 2 paper; the Llama 3 models are covered by Dubey et al. (2024), so the citations should be aligned with the models actually served.","section":"References and Section 2.1"},{"comment":"The example prompt contains empty <Reference></Reference> tags, which appears to be a rendering artifact and should either be filled in or removed. In addition, the phrase 'hoi polloi' in Section 1 is too informal for a journal venue.","section":"Appendix B.4"}],"recommendation":"major_revision","confidential_remarks":"The pilot deployment data are the genuine strength of this submission, and the authors are transparent about several limitations, which I weighed in favor of a major revision rather than rejection. The main risk is that the paper's headline claims (first-of-its-kind platform, lower-cost egalitarian gateway, reduced hallucination, HIPAA-grade privacy) are stated as contributions but not evidenced; the cost absence in particular is central to the value proposition. The editor may also wish to have the novelty claim checked, since 'first platform to address academic and research needs' is a strong first-to-market assertion that is hard to verify and is only weakly grounded in the short related-work section. Finally, the manuscript's writing quality is rough in places, with informal phrasing, duplicated appendix entries, and citation mismatches, so a thorough editorial pass is needed even after the substantive revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper is a straightforward systems description of AI-VERDE, an LLM gateway for universities that combines vLLM, LiteLLM, Weaviate, CILogon, and Kubernetes into a managed multi-tenant service with course-level access control and budget management. What is actually new is the integration: I don't know of another published account of a university-run platform that gives instructors automated API key provisioning, per-course RAG, and proxy access to commercial models with budget limits. The authors deployed it for real at the University of Arizona and report concrete usage numbers over six months: 78 users, 97,658 API calls, 110.68M tokens. That's credible engineering evidence of adoption, and the architecture section is clear enough for another institution to replicate.\n\nThe paper does several things well. The feature comparison table is honest, though it marks \"Yes\" for their own guardrails without detail. The discussion of institutional barriers—financial compliance, FERPA/HIPAA, onboarding overhead—reflects genuine engagement with campus IT realities. The appendix describes their survey honestly, including snowball sampling, which is a limitation they don't hide. The ethics statement is candid about proxy privacy limits and hardware constraints.\n\nThe soft spots are real, and the stress-test note is on target. The introduction promises a comparative cost study with commercial options, but none appears. There is no hardware cost, no personnel cost, no per-token or per-user cost, no baseline against OpenAI or Anvil. Table 2 shows 109.76M self-hosted tokens vs 0.919M proxy tokens, which suggests most traffic is on their own GPUs, but without fixed costs we cannot evaluate the \"much lower cost\" claim. This is load-bearing because the paper's central value proposition is egalitarian access, and egalitarian access rests on affordability. Second, the \"first of its kind\" claim is asserted, not verified by a literature search. Third, claims about privacy and reduced hallucination are plausible but unmeasured; the privacy claim is partially walked back in the ethics section. These are gaps a competent referee could push them to fix, not fatal flaws.\n\nVerdict: This is a useful infrastructure case study, not a research discovery. For a systems or demo track, it deserves a serious referee, but it needs major revision: add cost data, soften or substantiate the \"first\" claim, and report any security or accuracy evaluation. I'd bring it to a reading group only if the conversation is about LLM platform deployment in academia.\n\nRecommendation: send to peer review, conditional on major revisions.","headline":"Useful, honest systems paper about a University of Arizona LLM gateway; the cost comparison it promises never materializes, undercutting the central 'egalitarian' claim.","tokens_in":14869,"tokens_out":2036,"would_cite":false,"duration_ms":20879,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A university-run LLM gateway, tested for six months, handled 97,658 API calls and 110.7 million tokens for 78 users, and the paper argues this shows a practical route to campus-wide equitable access.","keywords":["LLM platform as a service","higher education","retrieval-augmented generation","privacy-preserving AI","API access management","on-premises model hosting","cost-effective AI access","university pilot study"],"falsifier":"Total the hardware, power, staff, and software cost of the pilot's 110.68 million self-hosted tokens and compare that with commercial API pricing for the same volume; if self-hosting costs as much or more, the cost-egalitarian claim fails.","tokens_in":13941,"feed_emoji":"🎓","tokens_out":9682,"duration_ms":81161,"temperature":0.7,"pith_summary":"AI-VERDE is a university-hosted gateway that lets instructors, students, and researchers reach both open-source and commercial language models through one login and one API key, with per-course budgets, access controls, and built-in retrieval-augmented generation. The paper claims this is the first platform designed to serve instructional and research needs together inside a higher-education institution, and that it removes the cost, privacy, and technical barriers that keep LLMs off campus. In a six-month pilot at a large public university, 78 users across five courses and ten research projects made 97,658 API calls worth 110.68 million tokens, mostly through self-hosted models. The sympathetic reading of the pilot is that even this modest voluntary adoption demonstrates a realistic route to campus-wide equitable access if hardware and staff support scale with demand.","feed_headline":"A campus AI gateway carried 110.7M tokens for 78 users","feed_subtitle":"Five courses and ten research projects shared one login, one API key, and one privacy boundary for six months.","key_machinery":"The carrying mechanism is a multi-tenant LLM gateway built around an OpenAI-compatible reverse-proxy API layer. Requests arrive at one access point and are routed to whatever model the caller names, whether a model hosted on institutional GPUs by a high-throughput serving engine or a commercial endpoint reached through a budget-controlled proxy. Around that core sit per-course vector databases for retrieval-augmented generation, a document-ingestion service, API-key and budget management, and single sign-on through the university's identity system. This stack converts the usual adoption problems, privacy, cost, authentication, model choice, and RAG setup, into configuration choices rather than per-user technical projects.","core_discovery":"The central claim is that one platform can unify all LLM access for an academic community and make that access egalitarian rather than dependent on personal subscriptions or technical skill. AI-VERDE is the proposed mechanism: each course or research group gets its own workspace, vector database, budget, and API keys, while all data stays on institutional infrastructure. The pilot numbers are offered as evidence: 97,658 API calls and 110.68 million tokens across 78 users, five courses, and ten research projects over six months. The paper reads this usage as \"significant engagement\" and \"huge potential for adoption,\" while acknowledging the raw numbers are low for an institution of that scale.","pith_inferences":["Going beyond what the paper proves, the metering data could be used to decide where to invest in GPUs versus commercial credits, since 99 percent of the pilot's tokens were self-hosted.","An unmeasured variable is the human support bundled into the platform; comparing adoption across departments with and without AI-specialist consultants would separate the software's contribution from the service's.","A direct extension would compare learning outcomes from the sandboxed course chatbot against a general-purpose chatbot, putting the hallucination-reduction claim on empirical footing."],"forward_implications":["A course can get a spending cap, automatically generated API keys for every enrolled student, and instant revocation at the end of the term.","Sensitive course and research data can be used with LLMs inside institutional infrastructure, which is what makes privacy-regulated work feasible.","Because the API matches a widely used standard, existing tools and code assistants work without modifications.","One gateway can serve instruction, research, and support departments, replacing several separate pay-per-use arrangements."],"supporting_citations":[{"why":"Supplies the high-throughput model-serving engine the platform uses to host open models.","marker":"Kwon et al., 2023"},{"why":"Defines retrieval-augmented generation, the RAG pipeline AI-VERDE wraps into per-course vector databases.","marker":"Lewis et al., 2020"},{"why":"Provides the Llama 3 family of open models served as a self-hosted option.","marker":"Dubey et al., 2024"},{"why":"Supplies Mistral 7B, one of the open models the gateway exposes.","marker":"Jiang et al., 2023b"},{"why":"Supplies the Phi-3 model family hosted by the platform.","marker":"Abdin et al., 2024"},{"why":"Describes the Anvil academic LLM service, the comparison baseline the paper measures AI-VERDE against.","marker":"Song et al., 2022"},{"why":"Documents the cyberinfrastructure partner that hosts the deployment and supplies compute.","marker":"Swetnam et al., 2024"},{"why":"Documents the NSF cloud resource the platform uses for scaling research workloads.","marker":"Hancock et al., 2021"},{"why":"Surveys LLM security and privacy risks that motivate the on-premises design.","marker":"Yao et al., 2024"},{"why":"Commercial ChatGPT is the cost and privacy baseline the platform positions itself against.","marker":"OpenAI, 2022"}],"fun_headline_variants":["One gateway, equal AI access for all campus groups","AI-VERDE unifies LLM access for courses and research","110M tokens, 78 users, one egalitarian platform","Campus AI hub levels LLM access for every team","From commercial to open source, one gateway for education"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that 78 voluntary pilot users and 372 survey respondents represent what the whole campus will need, and that the institution will continue to fund the hardware and human consultants as the user base grows.","fun_headline_variants_meta":{"raw":{"variants":["One gateway, equal AI access for all campus groups","AI-VERDE unifies LLM access for courses and research","110M tokens, 78 users, one egalitarian platform","Campus AI hub levels LLM access for every team","From commercial to open source, one gateway for education"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000468,"raw_usage":{"total_tokens":2275,"prompt_tokens":833,"completion_tokens":1442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":1360}},"tokens_in":449,"tokens_out":1442,"duration_ms":11869,"temperature":1.0,"reasoning_tokens":1360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:41:17.118017+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Total the hardware, power, staff, and software cost of the pilot's 110.68 million self-hosted tokens and compare that with commercial API pricing for the same volume; if self-hosting costs as much or more, the cost-egalitarian claim fails.","supporting_citations":[],"review_version":1}