Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models

T0 review · 2 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper assembles the first longitudinal dataset of public LLM-service outages and incident reports — eight services from OpenAI, Anthropic, and Character.AI — and uses it to show that ChatGPT failures take longer to resolve but occur…

desk verdict Useful dataset and descriptive stats, but the failure-isolation headline overstates what self-reported status data can show. read the letter →

arxiv 2501.12469 v2 pith:CPKTPQ67 submitted 2025-01-21 cs.PF cs.DC

classification cs.PFcs.DC
keywords FailurecharacterizationLLMfailure-recoveryreliabilityOpenAIAnthropicCharacter.AIoperationaldataanalytics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Public LLM services now carry hundreds of millions of users, yet no peer-reviewed, longitudinal record existed of how those services fail. This paper tries to fill that gap by scraping every outage and incident report the three major providers disclosed on their public status pages, covering eight services from early 2021 through August 2024, and by organizing those reports into one failure-recovery model. It then claims three headline facts: ChatGPT failures take longer to resolve than Claude's but happen less often; OpenAI and Anthropic incidents follow strong weekly and monthly periodicities; and OpenAI's services fail with better isolation than Anthropic's. The value of establishing these facts is that reliability engineering for LLM-dependent systems can move from guesswork to measured parameters.

What carries the argument

The paper's analytical engine is a five-status failure-recovery model distilled from industry status pages: Investigating (S1), Identified (S2), Monitoring (S3), Resolved (S4), and Postmortem (S5), which delimit Investigating, Repairing, Checking, and Learning periods. From these it derives Time To Resolve (S1 to S4), Time Between Failures (S1 to the next S1), and daily availability from scaled outage minutes, defined as major outage minutes plus 30 percent of partial outage minutes. The status markers let the paper compare recovery speeds across services, the period definitions let it say where recovery time is spent, and the scaled outage metric is what its availability rankings rest on.

What would settle it

Compare the paper's incident histories against timestamps reconstructed from public web archives of the same status pages, and against independent user-visible signals such as DownDetector-style reports or client-side error logs; if one provider's incidents are systematically missing, offset in start time, or resolved later than users observe, the MTTR, MTBF, periodicity, and isolation comparisons would need revision.

Watch

Extended reading notes

Core claim

On its own terms, the paper discovers that operator-reported LLM failures are measurable, patterned, and unevenly distributed. ChatGPT's median MTTR is 1.32 hours versus Claude's 0.93 hours, while Claude's median MTBF is 1.74 days versus ChatGPT's 2.04 days — so the two market-leading chatbots sit on opposite sides of the speed-versus-frequency tradeoff. Incident counts autocorrelate significantly up to weekly lag 12 for OpenAI and lag 7 for Anthropic, with weekday and work-hour peaks, supporting the periodicity claim. Same-day co-occurrence among Anthropic services exceeds 80 percent while cross-provider co-occurrence is near zero, supporting the isolation claim. The paper is explicit that these are properties of the self-reported status data, not root-cause explanations.

Load-bearing premise

The load-bearing premise is that the public status pages and incident reports self-disclosed by OpenAI, Anthropic, and Character.AI fully and consistently capture the failures under study; the paper itself acknowledges in its limitations section that the dataset is limited to what the operators report.

Editorial extensions

If this is right

  • Users of LLM services should expect at least one reportable incident per month for every service studied, so retries, caches, and fallback paths should be part of normal operation rather than emergency procedures.
  • Because Anthropic's three services fail together on most incident days and cross-provider co-occurrence is near zero, relying on a second provider is a viable high-availability strategy.
  • Because the checking period dominates resolution time, up to 82.71 percent of MTTR for Character.AI, faster validation and deployment pipelines could cut outage durations more than faster diagnosis.
  • The strong weekly and monthly autocorrelation of incident counts implies that failure frequency is partially predictable from its own history, opening the door to schedule-aware maintenance and user notification.
  • Despite high average availability, days with any outage are common — ChatGPT was fully available on only 88.85 percent of days — so availability averaged over months overstates the experience of a user who happens to hit a bad day.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper treats its dataset as complete for the three providers, but if a provider under-reports minor incidents or labels them inconsistently, the MTBF and co-occurrence comparisons could shift; that is an editorial judgment, not a finding of the paper.
  • The same status-page scraping and five-status modeling could be applied to other providers such as Gemini, Mistral, or DeepSeek to test whether the OpenAI and Anthropic patterns are general; the paper lists these as future work.
  • A natural next analysis is to align operator-reported incidents with independent signals such as DownDetector-style user reports or client-side error logging to check for reporting bias, a confirmation the paper itself flags as needed.
  • If the periodicity and isolation results hold, they suggest concrete service-level objectives: schedule risky changes outside weekday peaks and hold providers to isolation guarantees, since cross-provider outages are the rare case in this data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper presents a longitudinal dataset of outages and incidents from the public status pages of OpenAI, Anthropic, and Character.AI, covering eight LLM services over periods that begin as early as February 2021 and extend to August 2024. The authors introduce a failure-recovery model with status markers S1–S5 and periods (PI, PR, PC, PL), collect the data via automated scraping, and analyze recovery times (MTTR), inter-failure times (MTBF), temporal patterns, co-occurrence, and incident impact ranges. They report more than ten observations, including three headline claims: ChatGPT failures take longer to resolve but are less frequent than Claude's; OpenAI and Anthropic failures exhibit weekly and monthly periodicity; and OpenAI services offer better failure-isolation than Anthropic. Datasets and code are released on Zenodo and GitHub.

Significance. If the findings hold, this is one of the first comprehensive longitudinal empirical characterizations of public LLM service reliability, filling a clear gap in the dependability literature. The open release of curated datasets and code is a strong contribution that enables reproducibility and future research. The paper provides practically useful baseline statistics such as MTTR, MTBF, and daily availability for eight widely used services. The main interpretive risk is the causal framing of co-occurrence as 'failure isolation', which goes beyond what the status-page data can establish.

major comments (2)
  1. [Section 6.1, Eq. (4)] Equation (4) defines the conditional probability as P(A|B) = P(A∩B)/(P(A)P(B)) = O_AB/O_B. As printed, this is mathematically incorrect: the denominator contains an extra P(A) factor, and the correct expression is P(A|B) = P(A∩B)/P(B). The second equality to O_AB/O_B is the correct sample estimate of P(A|B), and the values in Figure 10b appear to use that definition, but the printed formula is wrong and should be corrected to avoid confusion in a paper whose central reproducibility argument relies on precise metric definitions.
  2. [Section 6.1 and Abstract, Observation #11] The abstract's claim that 'OpenAI services offer better failure-isolation than Anthropic services' is not supported by the evidence presented. The co-occurrence metric O_AB/O_B measures only whether two services appear in the operator's status reports on the same day; it cannot distinguish true architectural isolation from the provider's incident-attribution practice. This concern is magnified by the paper's own Figure 11, which shows that 71.79% of Anthropic incidents affect all three services versus 8.20% for OpenAI—the pattern expected if Anthropic commonly opens one incident and tags all services. Section 7 acknowledges general self-report bias, but the abstract and Observation #11 state the causal interpretation without that caveat. The authors should either rephrase the claim as 'co-occurrence in operator reports' or provide independent evidence (e.g., user-reported outage data or incident-ID structure) to separate reporting practice from actual isolation.
minor comments (6)
  1. [Section 3.2] The text states that OpenAI started reporting on February 11, 2023, but Table 3 lists OpenAI API and Playground start dates of 2021-02-11 and 2021-03-31, respectively; please reconcile this description with the table and with the abstract's '2021 to August 2024' range.
  2. [Section 4, Observation #7] There is a typo: 'Anthorpic' should be 'Anthropic', and 'small proposition of failures' should be 'small proportion of failures'.
  3. [Section 4] In the introductory paragraph of Section 4, 'Statues count' should be 'Statuses count'.
  4. [Section 6.1, Figure 10b note] The sentence 'The probability that service A is also outage while service B is outage' is grammatically awkward; consider rewording to 'The probability that service A also experiences an outage when service B experiences an outage'.
  5. [Figure 5b and Observation #6] Figure 5b shows the median MTBF for DALL·E as 4.63 days, while Observation #6 states 4.53 days; please align the text with the figure.
  6. [Section 3.2] The phrase 'out of the over 500 hundreds of incidents' is redundant; use 'over 500 incidents'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's observations are direct descriptive summaries of the scraped operator status data, with no fitted parameter presented as a prediction and no load-bearing self-citation chain.

full rationale

This is a descriptive observational study, not a derivation. The headline findings are arithmetic summaries of data collected from public status and incident pages (Section 3.2); no parameter is fitted to one subset and then 'predicted' for a closely related quantity. The failure-recovery model in Section 2.1 is presented as an industry convention rather than as a result derived from the data, and the 0.3 partial-outage weighting in Eq. (2) is an operator convention (ref [8]) applied uniformly, not a fitted coefficient. Metrics such as MTTR, MTBF, and the co-occurrence probability P(A|B)=O_AB/O_B in Eq. (4) are computed directly from the collected reports. Observation #12's 71.79% vs 8.20% values are direct counts of how providers tagged services in incident reports (Figure 11), and the isolation claim is an interpretation hedged in the text ('may be caused by a lack of isolation'). Section 7 explicitly acknowledges that the dataset is limited to what operators self-report; that is a real external-validity threat to the isolation interpretation, but it is a data-source limitation, not a circular step in which an output is equivalent to an input by construction. Self-citations ([13], [14], [53]) appear only in related-work comparisons and are not load-bearing for the central observations. The paper's claims are therefore self-contained descriptive statements about the scraped data, with no prediction that reduces to its inputs.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No new entities, forces, or dimensions are posited, and no statistical model is fitted, so the ledger contains no free parameters. The central measurements depend on five domain assumptions, chiefly the reliability of operator self-reports and the arbitrary 0.3 weighting of partial outages. These are stated or acknowledged in the paper, which is why the soundness score is moderate rather than low.

assumptions (5)
  • domain assumption Public status pages and incident reports from OpenAI, Anthropic, and Character.AI accurately capture the outages and incidents studied.
    The entire dataset is scraped from operator self-reports; Section 7 acknowledges the data are limited to what operators disclose. If operators omit or mislabel incidents, all comparative observations inherit the bias.
  • domain assumption Partial outages are weighted as 0.3 times major outages when computing daily availability.
    Equation (2) defines T_S = T_M + 0.3*T_P following Atlassian and OpenAI practice. This scaling is not derived or validated in the paper, but it directly shapes the availability statistics in Section 5.3 and Table 7.
  • ad hoc to paper Excluding 5 incidents whose status markers are out of order and ignoring the 'update' status does not bias the recovery metrics.
    Section 3.2 states these corner cases are excluded because none involves unusual durations; the paper provides no quantitative comparison with and without exclusion, so the neutrality of the exclusion is assumed.
  • domain assumption Incident timestamps reported in PDT are comparable across providers and across the study period.
    Figure 7 and Section 5.1 use PDT local times from status pages as if they are directly comparable for diurnal and weekly pattern analysis, with no adjustment for time zones or daylight-saving shifts in user behavior.
  • domain assumption The selected 8 services from 3 providers are representative of public LLM services generally.
    Section 3.1 uses data availability, popularity, and diversity as selection criteria; Section 7 notes that other service types (e.g., self-hosted or Google Gemini) are outside scope. The generality of the comparative claims rests on this selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models." pith.science (2026). https://pith.science/paper/CPKTPQ67

@misc{pith2026250112469,
  author       = {Pith},
  title        = {Pith review of: An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CPKTPQ67}},
  note         = {Machine review of arXiv:2501.12469}
}
read the original abstract

People and businesses increasingly rely on public LLM services, such as ChatGPT, DALLE, and Claude. Understanding their outages, and particularly measuring their failure-recovery processes, is becoming a stringent problem. However, only limited studies exist in this emerging area. Addressing this problem, in this work we conduct an empirical characterization of outages and failure-recovery in public LLM services. We collect and prepare datasets for 8 commonly used LLM services across 3 major LLM providers, including market-leads OpenAI and Anthropic. We conduct a detailed analysis of failure recovery statistical properties, temporal patterns, co-occurrence, and the impact range of outage-causing incidents. We make over 10 observations, among which: (1) Failures in OpenAI's ChatGPT take longer to resolve but occur less frequently than those in Anthropic's Claude;(2) OpenAI and Anthropic service failures exhibit strong weekly and monthly periodicity; and (3) OpenAI services offer better failure-isolation than Anthropic services. Our research explains LLM failure characteristics and thus enables optimization in building and using LLM systems. FAIR data and code are publicly available on https://zenodo.org/records/14018219 and https://github.com/atlarge-research/llm-service-analysis.

Figures

Figures reproduced from arXiv: 2501.12469 by the authors.

Figure 1
Figure 1. Monthly website visits, outages, and incidents for [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Visualization of the failure-recovery model with [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. Percent of time spent in the Investigating, Repair [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (8 more)
Figure 3
Figure 3. Figure 3: Presence of different status combinations, by ser [PITH_FULL_IMAGE:figures/full_fig_p005_3.png]
Figure 5
Figure 5. Figure 5: Distribution of MTTR and MTBF by service, with median values indicated. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: MTTR and MTBF by provider, ECDF plot. The closer [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Temporal distributions for incidents, PDT time. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Auto-correlations with the numbers of incidents [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Service daily availability by scaled outage minutes [%]. Some services started reporting later (see Section 3.2). [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Co-occurrence of outages between service pairs. [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Impact Services of OpenAI and Anthropic inci [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics

    cs.RO 2025-11 reject novelty 4.0 of 10

    Storing operator-validated diagnoses in a vector database and retrieving them during anomaly characterization cuts diagnostic dialog turns by 71% and raises characterization specificity from 2.7 to 4.8 in a BlueROV2 t...

Reference graph

Works this paper leans on

70 extracted references · 46 canonical work pages · cited by 1 Pith paper

  1. [1]

    Azure OpenAI Service

    2025. Azure OpenAI Service. https://azure.microsoft.com/en-us/products/ai- services/openai-service

  2. [2]

    Scaling Character.AI

    2025. Scaling Character.AI. https://cloud.google.com/blog/products/databases/ why-characterai-chose-spanner-and-alloydb-for-postgresql

  3. [3]

    Use Anthropic’s Claude models

    2025. Use Anthropic’s Claude models. https://cloud.google.com/vertex-ai/ generative-ai/docs/partner-models/use-claude

  4. [4]

    Jacob Alter, Ji Xue, Alma Dimnaku, and Evgenia Smirni. 2019. SSD failures in the field: symptoms, causes, and prediction models. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2019, Denver, Colorado, USA, November 17-19, 2019 , Michela Taufer, Pavan Balaji, and Antonio J. Peña (Eds.). ACM...

  5. [5]

    Anthropic. 2024. Release notes of Anthropic API. https://docs.anthropic.com/ en/release-notes/api

  6. [6]

    Anthropic. 2024. Release notes of Claude. https://docs.anthropic.com/en/release- notes/claude-apps

  7. [7]

    Internet Archive. 2024. Archived page of OpenAI status at 2024/04/10. https://web. archive.org/web/20240410235249/https://downdetector.com/status/openai/, Ac- cessed: 2024-09-30

  8. [8]

    Atlassian. 2024. Atlassian Support. https://support.atlassian.com/statuspage/ docs/display-historical-uptime-of-components/

Show all 70 references
  1. [9]

    Landwehr

    Algirdas Avizienis, Jean-Claude Laprie, Brian Randell, and Carl E. Landwehr

  2. [10]

    Ammar Ahmad Awan, Arpan Jain, Ching-Hsiang Chu, Hari Subramoni, and Dhabaleswar K. Panda. 2020. Communication Profiling and Characterization of Deep-Learning Workloads on Clusters With High-Performance Interconnects. IEEE Micro 40, 1 (2020), 35–43. doi:10.1109/MM.2019.2949986

  3. [11]

    Creager, Misha Efimov, Ilya Grigorik, Ben Jones, Harsha V

    Sam Burnett, Lily Chen, Douglas A. Creager, Misha Efimov, Ilya Grigorik, Ben Jones, Harsha V. Madhyastha, Pavlos Papageorge, Brian Rogan, Charles Stahl, and Julia Tuttle. 2020. Network Error Logging: Client-side measurement of end- to-end web service reliability. In 17th USENI...

  4. [12]

    Steven Wei Der Chien, Stefano Markidis, Chaitanya Prasad Sishtla, Luís Santos, Pawel Andrzej Herman, Sai Narasimhamurthy, and Erwin Laure. 2018. Charac- terizing Deep-Learning I/O Workloads in TensorFlow. In3rd IEEE/ACM Interna- tional Workshop on Parallel Data Storage & Data ...

  5. [13]

    Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Talluri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, and Alexandru Iosup. 2024. Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis. arXiv:2409.08949...

  6. [14]

    Xiaoyu Chu, Sacheendra Talluri, Laurens Versluis, and Alexandru Iosup. 2023. How Do ML Jobs Fail in Datacenters? Analysis of a Long-Term Dataset from an HPC Cluster. In Proceedings of the International Conference on Performance Engineering, Coimbra, Portugal, April, 2023

  7. [15]

    DownDetector. 2024. OpenAI user reports. https://downdetector.com/status/ openai/

  8. [16]

    Katz, Hemanth Kolla, Jacqueline Chen, Scott Klasky, and Manish Parashar

    Marc Gamell, Daniel S. Katz, Hemanth Kolla, Jacqueline Chen, Scott Klasky, and Manish Parashar. 2014. Exploring Automatic, Online Failure Recovery for Scientific Applications at Extreme Scales. In International Conference for High Performance Computing, Networking, Storage and...

  9. [17]

    Rohan Garg, Tirthak Patel, Gene Cooperman, and Devesh Tiwari. 2018. Shiraz: Ex- ploiting System Reliability and Application Resilience Characteristics to Improve Large Scale System Throughput. In48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks,...

  10. [18]

    Peter Garraghan, Paul Townend, and Jie Xu. 2014. An Empirical Failure-Analysis of a Large-Scale Cloud Computing Environment. In 15th International IEEE Sym- posium on High-Assurance Systems Engineering, HASE 2014, Miami Beach, FL, USA, January 9-11, 2014. IEEE Computer Society...

  11. [19]

    Google. 2019. Google Cluster Traces 2019. https://github.com/google/cluster- data?tab=readme-ov-file

  12. [21]

    Hochschild, Paul Turner, Jeffrey C

    Peter H. Hochschild, Paul Turner, Jeffrey C. Mogul, Rama Govindaraju, Parthasarathy Ranganathan, David E. Culler, and Amin Vahdat. 2021. Cores that don’t count. In HotOS ’21: Workshop on Hot Topics in Operating Systems, Ann Arbor, Michigan, USA, June, 1-3, 2021 , Sebastian Ang...

  13. [22]

    Williams

    Jiyao Hu, Zhenyu Zhou, Xiaowei Yang, Jacob Malone, and Jonathan W. Williams

  14. [23]

    Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen, and Tianwei Zhang. 2021. Characterization and prediction of deep learning workloads in large-scale GPU datacenters. In International Conference for High Performance Computing, Net- working, Storage and Analysis, SC 2021, St. Lou...

  15. [24]

    Qinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang, Meng Zhang, Qiaoling Chen, Peng Sun, Dahua Lin, Xiaolin Wang, Yingwei Luo, Yonggang Wen, and Tianwei Zhang. 2024. Characterization of Large Language Model Development in the Datacenter. In 21st USENIX Symposium on Networked Sy...

  16. [25]

    Singularity Hub. 2022. OpenAI Says DALL-E Is Generating Over 2 Million Images a Day. https://singularityhub.com/2022/10/03/openai-says-dall-e-is-generating- over-2-million-images-a-day-and-thats-just-table-stakes/

  17. [26]

    IDC. 2024. A Deep Dive Into Global AI and Generative AI Spend- ing. https://blogs.idc.com/2024/08/16/a-deep-dive-into-idcs-global-ai-and- generative-ai-spending/

  18. [27]

    Anthropic Incident. 2024. https://status.anthropic.com/history

  19. [28]

    Character.AI Incident. 2024. https://status.character.ai/history

  20. [29]

    OpenAI Incident. 2024. https://status.openai.com/history

  21. [30]

    Bahman Javadi, Derrick Kondo, Alexandru Iosup, and Dick H. J. Epema. 2013. The Failure Trace Archive: Enabling the comparison of failure measurements and models of distributed systems. J. Parallel Distributed Comput. 73, 8 (2013), 1208–1223. doi:10.1016/J.JPDC.2013.04.002

  22. [31]

    Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In Proceedings of the 2019 USENIX Annual Technical Conference, USENIX ATC 2019, Renton, W A, US...

  23. [32]

    Jing Yu Koh, Daniel Fried, and Russ Salakhutdinov. 2023. Generating Im- ages with Multimodal Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December ...

  24. [33]

    Małgorzata Łazuka, Andreea Anghel, and Thomas Parnell. 2024. LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services. arXiv preprint arXiv:2410.02425 (2024)

  25. [34]

    Baolin Li, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle, Matthew Hubbell, Michael Jones, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Jul...

  26. [35]

    Sullivan, Timothy Tsai, Karthik Pattabiraman, Joel S

    Guanpeng Li, Siva Kumar Sastry Hari, Michael B. Sullivan, Timothy Tsai, Karthik Pattabiraman, Joel S. Emer, and Stephen W. Keckler. 2017. Understanding error propagation in deep learning neural network (DNN) accelerators and applications. In Proceedings of the International Co...

  27. [36]

    Jie Li, Rui Wang, Ghazanfar Ali, Tommy Dang, Alan Sill, and Yong Chen. 2023. Workload Failure Prediction for Data Centers. In 16th IEEE International Confer- ence on Cloud Computing, CLOUD 2023, Chicago, IL, USA, July 2-8, 2023 . IEEE, 479–485. doi:10.1109/CLOUD60044.2023.00064

  28. [37]

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural Infor- mation Processing Systems 36: Annual Conference on Neural Info...

  29. [38]

    Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang

  30. [39]

    Kalbarczyk, Ravishankar K

    Catello Di Martino, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Fabio Baccanico, Joseph Fullop, and William Kramer. 2014. Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue Waters. In 44th Annual IEEE/IFIP International Conference on Dependabl...

  31. [40]

    OpenAI. 2024. Elevated errors in ChatGPT. https://status.openai.com/incidents/ w20mcckg1748, Accessed: 2024-09-30

  32. [41]

    OpenAI. 2024. Start using ChatGPT instantly (Apr 1, 2024). https://help.openai. com/en/articles/6825453-chatgpt-release-notes

  33. [42]

    Rahul Potharaju and Navendu Jain. 2013. When the network crumbles: an empirical study of cloud network failures and their impact on services. In ACM Symposium on Cloud Computing, SOCC ’13, Santa Clara, CA, USA, October 1-3, 2013, Guy M. Lohman (Ed.). ACM, 15:1–15:17. doi:10.11...

  34. [43]

    Argyraki, and Edouard Bugnion

    Mia Primorac, Katerina J. Argyraki, and Edouard Bugnion. 2021. When to Hedge in Interactive Services. In 18th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2021, April 12-14, 2021, James Mickens and Renata Teix- eira (Eds.). USENIX Association, 373–387....

  35. [44]

    Similarweb Pro. 2024. https://pro.similarweb.com/

  36. [45]

    Ganger, Randy H

    Charles Reiss, Alexey Tumanov, Gregory R. Ganger, Randy H. Katz, and Michael A. Kozuch. 2012. Heterogeneity and dynamicity of clouds at scale: Google trace analysis. In ACM Symposium on Cloud Computing, SOCC ’12, San Jose, CA, USA, October 14-17, 2012 , Michael J. Carey and St...

  37. [46]

    Reuters. 2023. ChatGPT sets record for fastest-growing user base - ana- lyst note. https://www.reuters.com/technology/chatgpt-sets-record-fastest- growing-user-base-analyst-note-2023-02-01/

  38. [47]

    Reuters. 2024. Google-backed Anthropic releases Claude chatbot across Eu- rope. https://www.reuters.com/technology/google-backed-anthropic-releases- claude-chatbot-across-europe-2024-05-13/

  39. [48]

    Reuters. 2024. OpenAI says ChatGPT’s weekly users have grown to 200 mil- lion. https://www.reuters.com/technology/artificial-intelligence/openai-says- chatgpts-weekly-users-have-grown-200-million-2024-08-29/

  40. [49]

    Chen, and Walter Binder

    Andrea Rosà, Lydia Y. Chen, and Walter Binder. 2015. Predicting and Mitigating Jobs Failures in Big Data Clusters. In 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2015, Shenzhen, China, May 4-7, 2015 . IEEE Computer Society, 221–230. doi:1...

  41. [50]

    Roumeliotis and Nikolaos D

    Konstantinos I. Roumeliotis and Nikolaos D. Tselikas. 2023. ChatGPT and Open- AI Models: A Preliminary Review. Future Internet 15, 6 (2023). doi:10.3390/ fi15060192

  42. [51]

    Selenium. 2024. WebDriver Documentation. https://www.selenium.dev/ documentation/webdriver/

  43. [52]

    Siqi Shen, Alexandru Iosup, Assaf Israel, Walfredo Cirne, Danny Raz, and Dick H. J. Epema. 2015. An Availability-on-Demand Mechanism for Datacenters. In 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2015, Shenzhen, China, May 4-7, 2015 . IE...

  44. [53]

    Abad, and Alexandru Iosup

    Sacheendra Talluri, Alicja Luszczak, Cristina L. Abad, and Alexandru Iosup. 2019. Characterization of a Big Data Storage Workload in the Cloud. In Proceedings of the 2019 ACM/SPEC International Conference on Performance Engineering, ICPE 2019, Mumbai, India, April 7-11, 2019 ,...

  45. [54]

    Sacheendra Talluri, Leon Overweel, Laurens Versluis, Animesh Trivedi, and Alexandru Iosup. 2021. Empirical Characterization of User Reports about Cloud Failures. In IEEE International Conference on Autonomic Computing and Self- Organizing Systems, ACSOS 2021, Washington, DC, U...

  46. [55]

    Rogers, Don Maxwell, Paolo Rech, Sudharshan S

    Devesh Tiwari, Saurabh Gupta, James H. Rogers, Don Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, and Arthur S. Bland. 2015. Understanding GPU errors on large-scale HPC systems and...

  47. [56]

    Anthropic Uptime. 2024. https://status.anthropic.com/uptime

  48. [57]

    Character.AI Uptime. 2024. https://status.character.ai/uptime

  49. [58]

    OpenAI Uptime. 2024. https://status.openai.com/uptime

  50. [59]

    Glassman

    Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In CHI ’22: CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April 2022...

  51. [60]

    Laurens Versluis, Mehmet Çetin, Caspar Greeven, Kristian Laursen, Damian Podareanu, Valeriu Codreanu, Alexandru Uta, and Alexandru Iosup. 2023. Less is not more: We need rich datasets to explore. Future Gener. Comput. Syst. 142 (2023), 117–130. doi:10.1016/J.FUTURE.2022.12.022

  52. [61]

    Jiayin Wang, Weizhi Ma, Peijie Sun, Min Zhang, and Jian-Yun Nie. 2024. Un- derstanding User Experience in Large Language Model Interactions. CoRR abs/2401.08329 (2024). doi:10.48550/ARXIV.2401.08329 arXiv:2401.08329

  53. [62]

    Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Zhenheng Tang, Xin He, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, and Xiaowen Chu. 2024. BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems. arXiv:2401.17644

  54. [63]

    Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the Wild: Workload Anal- ysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Im...

  55. [64]

    WIRED. 2024. The obsession with Character AI is becoming more common. https://wired.me/technology/character-ai-obsession/

  56. [65]

    Yuchen Xia, Jiho Kim, Yuhan Chen, Haojie Ye, Souvik Kundu, Cong Hao, and Nishil Talati. 2024. Understanding the Performance and Estimating the Cost of LLM Fine-Tuning. CoRR abs/2408.04693 (2024). doi:10.48550/ARXIV.2408.04693 arXiv:2408.04693

  57. [66]

    Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Ben Hu. 2024. Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond. ACM Trans. Knowl. Discov. Data 18, 6 (2024), 160:1–160:32. doi:10.1145/3649506

  58. [67]

    FOX5 New York. 2024. ChatGPT recovers following outage affecting thousands of users. https://www.fox5ny.com/news/chatgpt-down-users-openai-internet, Accessed: 2024-09-30

  59. [68]

    Yang Zhang, Bogdan Vasilescu, Huaimin Wang, and Vladimir Filkov. 2018. One size does not fit all: an empirical study of containerized continuous deployment workflows. In Proceedings of the 2018 ACM Joint Meeting on European Software Engineering Conference and Symposium on the ...

  60. [2004]

    IEEE Trans

    Basic Concepts and Taxonomy of Dependable and Secure Computing. IEEE Trans. Dependable Secur. Comput. 1, 1 (2004), 11–33. doi:10.1109/TDSC.2004.2

  61. [2020]

    In 17th USENIX Symposium on Networked Sys- tems Design and Implementation, NSDI 2020, Santa Clara, CA, USA, February 25-27, 2020, Ranjita Bhagwan and George Porter (Eds.)

    CableMon: Improving the Reliability of Cable Broadband Networks via Proactive Network Maintenance. In 17th USENIX Symposium on Networked Sys- tems Design and Implementation, NSDI 2020, Santa Clara, CA, USA, February 25-27, 2020, Ranjita Bhagwan and George Porter (Eds.). USENIX...

  62. [2023]

    LLMScore: Unveiling the Power of Large Language Models in Text- to-Image Synthesis Evaluation. In Advances in Neural Information Process- ing Systems 36: Annual Conference on Neural Information Processing Sys- tems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 20...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.