{"id":"fde9a518-ed5a-4180-8390-050f619bb9cd","arxiv_id":"2508.21377","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey applying the Kaddour et al. challenge taxonomy to GPT-4o and DeepSeek-V3-0324, concluding closed models favor safety while open models favor cost and customization.","lead":"This survey compares OpenAI's GPT-4o against open-source DeepSeek-V3-0324 across 16 LLM challenges and 8 application areas. It concludes that closed models win on safety and reliability while open models win on cost, flexibility, and control, so the best choice depends on the use case.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Safety/reliability verdicts rest on uncited quantitative claims; if these numbers fail verification, the central trade-off loses its empirical backbone.","rationale":"The reader's weakest_assumption correctly identifies the uncited and non-reproducible factual basis for the safety and reliability verdicts. My stress-test reaches the same conclusion: the central claim is a qualitative trade-off that is plausible and consistent with consensus, but the paper's own quantitative evidence for it is not independently verifiable. This does not overturn the qualitative guidance, but it does mean the paper cannot be accepted unconditionally as a reliable decision-making reference. The reader's CONDITIONAL verdict is appropriate; the condition should be full sourcing of all numerical claims and resolution of the parameter-count inconsistency. I do not see a more load-bearing concern than this: the paper is explicitly a survey, not a novel empirical study, and its high-level conclusions are robust to minor numerical corrections. The uncited numbers are the softest point because they are presented as objective measurements supporting specific model choices, yet they lack the transparency the paper itself demands of model documentation.","tokens_in":14690,"tokens_out":3106,"duration_ms":34179,"concrete_test":"Compile a full list of quantified claims in Sections II.B, III.H, and III.I (89% accuracy, 82% disallowed-content reduction, 3.9% vs 1.5% hallucination rates, 77% attack success, 35.6% WMD pass, 53.3% hate speech, 48.9% self-harm). Attempt to trace each to a primary source (OpenAI system card, DeepSeek-V3 technical report, Vectara HHEM leaderboard, or named independent audit). For any claim that cannot be sourced or that applies to a different model version/date, mark it unverified and recompute the Section III verdicts with only verified evidence. Then check whether the safety advantage of GPT-4o remains 'superior' or degrades to 'mixed/unknown'. Also resolve the 671B vs 685B parameter discrepancy by consulting the DeepSeek-V3 technical report.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the trade-off: GPT-4o for safety/reliability, DeepSeek for control/cost. The safety half of that trade-off is supported almost entirely by uncited or non-reproducible numbers. Section II.B states GPT-4o achieves 89% accuracy and reduces disallowed content by 82%, with no reference. Section III.H cites Vectara hallucination rates (3.9% vs 1.5%) without a source or measurement details. Section III.I invokes 'independent audits' reporting 77% attack success and 35.6% WMD pass rate, again without citation. These numbers are not incidental; they are the only quantified evidence that GPT-4o is meaningfully safer and more reliable than DeepSeek. The informal screenshot tests (Figures 3–7) are anecdotal and non-reproducible. The paper acknowledges both models are opaque (Sections III.O and III.P), yet it states these figures as settled facts. If any of these numbers are wrong, non-comparable, or cherry-picked, the conclusion that GPT-4o is the safer choice is unsupported. The parameter-count inconsistency (671B in Section I vs 685B in Section III.C) further signals that quantitative details are not carefully checked. The qualitative trade-off may survive because it aligns with field consensus, but the paper's specific evidentiary base for it is fragile.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a survey-style comparison of GPT-4o and DeepSeek-V3-0324 across 16 LLM challenges, organized into design, behavioral, and evaluation issues, followed by application-level recommendations. The central thesis is that GPT-4o is preferable when safety, reliability, and minimal maintenance are priorities, while DeepSeek is preferable when control, customization, and cost are priorities. The paper claims to 'showcase' these trade-offs using a combination of literature-based descriptions, uncited benchmark figures, and informal screenshot demonstrations.","tokens_in":14850,"tokens_out":6281,"duration_ms":64667,"significance":"The paper's contribution is mainly organizational and practical. It provides a readable taxonomy of known LLM challenges and maps them onto concrete model choices. The qualitative direction of the trade-off (closed models being safer and more polished, open models being more efficient and controllable) agrees with broad field consensus, and the paper cites some relevant sources, including the DeepSeek-V3 technical report and a recent medical-exam evaluation. However, the manuscript does not provide machine-checked proofs, code, or reproducible experiments. Its quantitative evidence is largely uncited, and the informal tests do not meet a reproducibility standard. As such, the central claim is not yet supported at the level the paper's confident language suggests.","major_comments":[{"comment":"Section II.B introduces four quantitative performance/safety claims (89% accuracy, 87% precision, 82% disallowed-content reduction, 2x faster inference, 50% lower cost) with no reference. These are the paper's first concrete evidence for the GPT-4o reliability advantage and they recur as implicit support in Sections III.A, III.H, and III.I. Please provide primary sources or remove the numbers; if removed, the subsequent safety verdicts must be re-grounded in citable evidence rather than assertion.","section":"Section II.B"},{"comment":"Section III.H states, 'According to Vectara's HHEM 2.1 benchmark, DeepSeek has a hallucination rate of 3.9%, compared to GPT-4o's 1.5%' and adds a 14.3% figure for DeepSeek-R1. No date, URL, test-set description, or configuration is given, and Sections III.O and III.P describe both models as opaque. Since the hallucination comparison is one of the two central pillars of the safety/reliability verdict, please supply the exact benchmark version, dataset, prompt settings, and access date, and indicate whether the rates are directly comparable.","section":"Section III.H"},{"comment":"Section III.I invokes 'independent audits' for a series of safety metrics (77% prompt-attack success, 69.2% evasion, 35.6% WMD pass, 53.3% hate speech, 48.9% self-harm, and failure on all Pliny injections) without citation. These numbers are the sole quantitative support for the conclusion that 'GPT-4o is clearly superior in safety and alignment.' Please cite the specific audits (organization, date, methodology, and model checkpoint), or label these as author-conducted stress tests with a full protocol. As written, the evidence is not verifiable.","section":"Section III.I"},{"comment":"Figures 3-7 are presented as results of 'our testing' with no protocol. There is no description of the number of prompts, selection criteria, model version/API parameters, temperature, or environment; the screenshots are therefore anecdotal and potentially unrepresentative. Since these examples are used to illustrate and partially justify the safety and tokenization verdicts, please either add a methodology appendix specifying how the prompts were generated and chosen, or explicitly downgrade the figures to illustrative, non-evidentiary examples.","section":"Figures 3-7 and Sections III.A, III.B, III.I"},{"comment":"The parameter count for DeepSeek is inconsistent across the manuscript: 671B in Sections I and II.C, but '685B' in Section III.C. The DeepSeek-V3 technical report gives 671B total with 37B active. This inconsistency, combined with the uncited numbers above, suggests that the quantitative details have not been carefully checked. Please correct the value and audit the remaining quantitative statements for accuracy.","section":"Section III.C vs. Sections I and II.C"}],"minor_comments":[{"comment":"Typo: 'close source' should be 'closed source.' The title 'GPT and DeepSeek family of models' is grammatically awkward; consider 'GPT and DeepSeek Families of Models.'","section":"Abstract and Title"},{"comment":"The statement that GPT-4o has 'an estimated several hundred billion parameters' is vague and unsourced; provide a citation or mark it as an author estimate.","section":"Section II.B"},{"comment":"The training-cost estimates ($5-6M for DeepSeek, >$100M for GPT-4) are presented without a source. Please cite the underlying reports and clarify that these are estimates.","section":"Section III.C"},{"comment":"The claim that 'users report steep performance drops' after ~20K tokens and failures at ~56K tokens is not cited. Add a source or explicitly label it as anecdotal.","section":"Section III.F"},{"comment":"The claim that DeepSeek's temperature setting of 1.0 corresponds to an effective temperature of 0.3 is not sourced; cite the DeepSeek documentation or remove the statement.","section":"Section III.G"},{"comment":"The statement that 'GPT-4o powers trusted tools like Duolingo and Khan Academy' needs a citation or should be hedged, as these integrations may change over time.","section":"Section IV.H"},{"comment":"Several statements in Section III (e.g., inverse scaling in III.N, watermarking in III.M) could be supported by existing references [8], [13], etc.; please add citations at the point of claim rather than only in the reference list.","section":"References"},{"comment":"The architecture diagrams are labeled 'illustrative' but no source is given. Please state whether they are original or redrawn from existing sources, and cite the source if applicable.","section":"Figures 1-2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as a practitioner-oriented industry report rather than a standard scholarly research paper. The central quantitative claims are presented with unusually little support; the uncited benchmark figures and the absence of a testing protocol are the main barriers to acceptance. I would require that all quantitative claims either be removed or properly cited, and that the informal tests either be given a protocol or be explicitly reframed as illustrative. The author affiliations are noted but are not, by themselves, a reason to reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a clearly written but derivative survey: it applies Kaddour et al.'s 16-challenge taxonomy to GPT-4o and DeepSeek-V3 and lands on the expected closed-versus-open trade-off. The qualitative conclusions are consistent with the field's consensus, and the application-by-application guidance is sensible. What is genuinely new is a small set of hand-run prompt comparisons (Figures 3–7) and a domain recommendation matrix; neither is rigorous, but they are clearly presented as illustrative, and the paper is honest about the opacity of both models (Sections III.O and III.P).\n\nThe soft spots are real but not fatal to the overall conclusion. Specific figures—the 89% accuracy, the 82% disallowed-content reduction, the Vectara hallucination rates (3.9% vs 1.5%), and the DeepSeek attack-success numbers (77%, 35.6%)—are all uncited. That matters because the safety verdict rests on them. The paper also mixes the parameter count for DeepSeek (671B in Section I, 685B in Section III.C), which suggests the quantitative details weren't checked carefully. The informal screenshots are anecdotes, not a protocol, and the paper does not pretend otherwise.\n\nMy take: the central trade-off claim is robust enough to survive these problems—closed models do tend to be safer and more reliable, open models more adaptable and cheaper—so the paper is not misleading readers. But it is a practitioner primer, not a research contribution. The taxonomy is inherited, the specific numbers are unverified, and the original content is minor. I would not cite it in my own work, and I would not expect a serious archival venue to send it out for full review unless the authors add sources for every quantitative claim and resolve the parameter-count contradiction. For a team actively choosing between GPT-4o and DeepSeek, it is a reasonable conversation starter and a broad map of issues, but I would not treat its numbers as authoritative.\n\nRecommendation: desk reject at a top-tier venue; accept with heavy revision at a practitioner-oriented outlet only if the sourcing is fixed. If you are on the fence, give it to a junior person as a quick orientation document.\n\nBest","headline":"A useful but derivative practitioner survey whose qualitative trade-offs are right and whose quantitative assertions need sourcing.","tokens_in":15456,"tokens_out":3385,"would_cite":false,"duration_ms":37424,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o and DeepSeek are complementary trade-offs, not ranked alternatives: closed models offer safety, open models offer control.","keywords":["large language models","GPT-4o","DeepSeek","closed source vs open source","model safety","hallucination","Mixture-of-Experts","model selection"],"falsifier":"Take one fixed evaluation battery: the same hallucination test set, the same long-context documents beyond 20K tokens, the same safety and prompt-injection probes, and the same math and coding benchmarks, and run it on GPT-4o and DeepSeek-V3-0324 under identical conditions. If DeepSeek's hallucination rate is not higher than GPT-4o's, or GPT-4o does not beat DeepSeek on safety probes, the paper's central trade-off claim would fail its empirical test.","tokens_in":14459,"feed_emoji":"🤖","tokens_out":6296,"duration_ms":60084,"temperature":0.7,"pith_summary":"This paper is a structured comparison of two contemporary large language models: the closed GPT-4o and the open-weight DeepSeek-V3-0324. It argues that neither model is simply better; instead, each embodies a trade-off. The closed model offers stronger safety alignment, lower hallucination, and more predictable behavior in high-stakes and user-facing settings, while the open model offers lower training and deployment cost, full customization, and transparency. The paper works through sixteen known LLM challenges and then maps the two models' relative strengths onto application domains, concluding that model choice should be driven by whether a use case prioritizes safety and reliability or control, privacy, and cost.","feed_headline":"Closed or open LLM? The choice is risk vs control","feed_subtitle":"A 16-challenge comparison maps GPT-4o's safety edge against DeepSeek's efficiency and customization edge.","key_machinery":"The paper's organizing instrument is a taxonomy of sixteen LLM challenges, grouped into design, behavioral, and evaluation categories and inherited from [5]. Each challenge is applied to GPT-4o and DeepSeek as a fixed comparison lens, yielding a per-challenge verdict; those verdicts are then mapped onto seven application domains to produce model recommendations. The recurring mechanism that carries the argument is the contrast between a closed, centrally aligned deployment, which delivers safety by design at the cost of transparency and control, and an open-weight, sparsely activated Mixture-of-Experts model, which delivers efficiency and adaptability by design at the cost of built-in safety","core_discovery":"The paper's central claim is that the practical difference between GPT-4o and DeepSeek-V3-0324 is best described as a closed-source versus open-source trade-off. GPT-4o is presented as the safer, more reliable, more polished system: reinforcement learning from human feedback, red-teaming, content filtering, a reported hallucination rate of 1.5% versus DeepSeek's 3.9%, reliable handling of long contexts, and strong refusal behavior make it the recommended default for chatbots, content creation, education, and high-stakes advice. DeepSeek is presented as the more efficient and adaptable system: a Mixture-of-Experts architecture with about 37B active parameters per token, roughly $5–6 million t","pith_inferences":["The comparison's empirical edge cases are only as strong as third-party numbers that the paper cites without sources; a reproducible benchmark battery would be the natural next step.","The closed-versus-open axis likely generalizes to other API-based aligned models and open-weight models, but the specific margin of safety versus flexibility will shift as new versions land.","A testable extension: run the same sixteen-challenge rubric on later model versions, or on fine-tuned open models with added alignment, to see whether the safety gap can be closed by community effort.","The paper implicitly assumes that informal screenshot probes represent typical usage; formalizing those probes into a fixed prompt set would make the qualitative findings falsifiable."],"forward_implications":["A use case can be screened by four questions: Is it user-facing? Is the cost of error high? Is the data confidential? Does the team need full customization? The answers point to GPT-4o or DeepSeek respectively.","Open models at DeepSeek's cost level put competitive pressure on closed providers, and the paper expects MoE-style efficiency techniques to be absorbed into next-generation closed models.","Hybrid deployments become viable: a closed, aligned model for the front end and an open model for internal document processing or specialized coding.","If the safety gap is real, deploying open models in public-facing or regulated settings requires an extra layer of monitoring and filtering that the paper treats as the adopter's responsibility.","The paper's verdicts imply that model selection guidance should be updated frequently, because both models are versioned and the trade-off magnitudes will shift."],"supporting_citations":[{"why":"Supplies the sixteen-challenge taxonomy that structures the entire comparison and the application landscape.","marker":"[5]"},{"why":"Technical report for DeepSeek-V3; provides the MoE architecture, MLA, MTP, training dataset size, and cost figures the efficiency verdicts rest on.","marker":"[11]"},{"why":"GPT-4 technical report; establishes the closed-model baseline that GPT-4o is described as improving with a larger context window and fine-tuning refinements.","marker":"[12]"},{"why":"The RLHF method paper; underwrites the alignment and safety mechanism credited to GPT-4o throughout the comparison.","marker":"[7]"},{"why":"Retrieval-augmented generation; supplies the mechanism both models use to address outdated knowledge in the relevant challenge section.","marker":"[10]"},{"why":"Head-to-head medical exam study comparing GPT-4o and DeepSeek-R1; used as evidence that an open model can match a closed one in a specialized domain.","marker":"[6]"},{"why":"Transformer architecture paper; foundational context for both models' design and for the distinction between dense and sparse MoE architectures.","marker":"[1]"}],"fun_headline_variants":["LLMs: safety or control—the core trade-off","GPT-4o vs DeepSeek: risk versus control","Open vs closed LLMs: pick your risk","LLM choice: robust safety or full control","Which LLM? Safer but rigid, or flexible and risky"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the performance and safety numbers quoted for both models, such as GPT-4o's 89% accuracy and DeepSeek's 3.9% hallucination rate, are accurate, current, and directly comparable; the paper does not provide primary sources for several of them, and its informal screenshot tests are assumed to represent typical usage.","fun_headline_variants_meta":{"raw":{"variants":["LLMs: safety or control—the core trade-off","GPT-4o vs DeepSeek: risk versus control","Open vs closed LLMs: pick your risk","LLM choice: robust safety or full control","Which LLM? Safer but rigid, or flexible and risky"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1374,"prompt_tokens":705,"completion_tokens":669,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":603}},"tokens_in":449,"tokens_out":669,"duration_ms":7489,"temperature":1.0,"reasoning_tokens":603,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:19:19.643489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one fixed evaluation battery: the same hallucination test set, the same long-context documents beyond 20K tokens, the same safety and prompt-injection probes, and the same math and coding benchmarks, and run it on GPT-4o and DeepSeek-V3-0324 under identical conditions. If DeepSeek's hallucination rate is not higher than GPT-4o's, or GPT-4o does not beat DeepSeek on safety probes, the paper's central trade-off claim would fail its empirical test.","supporting_citations":[{"cited_title":"Training Language Models to Follow Instructions with Human Feedback,","cited_arxiv_id":null,"evidence_quote":"The RLHF method paper; underwrites the alignment and safety mechanism credited to GPT-4o throughout the comparison."},{"cited_title":"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,","cited_arxiv_id":null,"evidence_quote":"Retrieval-augmented generation; supplies the mechanism both models use to address outdated knowledge in the relevant challenge section."},{"cited_title":"Performance of GPT-4o and DeepSeek-R1 in the Polish Infectious Diseases Specialty Exam,","cited_arxiv_id":null,"evidence_quote":"Head-to-head medical exam study comparing GPT-4o and DeepSeek-R1; used as evidence that an open model can match a closed one in a specialized domain."},{"cited_title":"Attention is All You Need,","cited_arxiv_id":null,"evidence_quote":"Transformer architecture paper; foundational context for both models' design and for the distinction between dense and sparse MoE architectures."}],"review_version":1}