Pith. sign in

REVIEW 4 cited by

Operationalizing Machine Learning: An Interview Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.09125 v1 pith:F32P7XGE submitted 2022-09-16 cs.SE cs.HCcs.LG

classification cs.SEcs.HCcs.LG
keywords productiondeploymentperformanceexperimentationimplicationsinterviewslearningmachine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Organizations rely on machine learning engineers (MLEs) to operationalize ML, i.e., deploy and maintain ML pipelines in production. The process of operationalizing ML, or MLOps, consists of a continual loop of (i) data collection and labeling, (ii) experimentation to improve ML performance, (iii) evaluation throughout a multi-staged deployment process, and (iv) monitoring of performance drops in production. When considered together, these responsibilities seem staggering -- how does anyone do MLOps, what are the unaddressed challenges, and what are the implications for tool builders? We conducted semi-structured ethnographic interviews with 18 MLEs working across many applications, including chatbots, autonomous vehicles, and finance. Our interviews expose three variables that govern success for a production ML deployment: Velocity, Validation, and Versioning. We summarize common practices for successful ML experimentation, deployment, and sustaining production performance. Finally, we discuss interviewees' pain points and anti-patterns, with implications for tool design.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs

    cs.SE 2025-09 unverdicted novelty 7.0 of 10

    The submission's abstract claims a CPG-based ML smell detector with 88.14% recall, but the full text implements an AST-only DSL detector with 88.89% recall, and no CPG experiment appears.

  2. Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software

    cs.ET 2026-08 conditional novelty 6.0 of 10

    A seven-layer, vendor-agnostic architecture using open-source tools to detect, diagnose, fix, verify, and remember pipeline failures, with humans approving risky actions.

  3. Better Training Data Attribution via Better Inverse Hessian-Vector Products

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ASTRA, an EKFAC-preconditioned Neumann series iteration, computes more accurate inverse Hessian-vector products and improves training data attribution scores over EKFAC baselines.

  4. FaaS and Furious: abstractions and differential caching for efficient data pre-processing

    cs.DB 2024-11 conditional novelty 5.0 of 10

    A columnar differential cache for lakehouse pipelines reuses overlapping scan fragments and reduces S3 bytes read by up to 30% in preliminary benchmarks.

Pith tools