REVIEW 1 major objections 7 minor 34 references
AI Strategy: How to Choose What AI Product to Implement
T0 review · 1 major / 7 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Coarse ratings of value-if-successful, likelihood of success, and investment are enough to separate strong AI product bets from weak ones before you build.
desk verdict A usable AI-project prioritization framework with honest limits: the Compass cases illustrate well but do not validate discriminative power. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
eROI (expected ROI): the proportional relation Value if Successful × Likelihood of Success / Investment Required, with each pillar broken into sub-dimensions, rated on a coarse Low-to-Very-High scale, Likelihood gated by its weakest certainty source, and projects judged by full profiles assembled into a mixed portfolio rather than collapsed into one ROI number.
What would settle it
Apply the same coarse three-pillar rating process, before outcomes are known, to a fresh slate of AI candidates at firms that were not involved in writing the framework; check whether high-eROI profiles systematically outperform low-eROI ones on realized business results, and whether the Time-on-Market-style scientific-certainty gate correctly predicts projects that later stall.
Extended reading notes
Core claim
Precise pre-build ROI estimates are usually infeasible for AI products; decomposing each bet into separately rated Value if Successful, Likelihood of Success (gated by product, technical, data, and scientific certainty), and Investment Required, then using coarse ordinal profiles plus portfolio assembly rather than a single top-ranked pick, is enough to tell strong AI product bets from weak ones and dissolves the catch-22 that teams cannot estimate ROI until they know whether the project will work.
Load-bearing premise
That the three author-involved, retrospective Compass cases—rated with outcome hindsight by people who were among the original decision-makers—are enough to show the framework can tell strong bets from weak ones in use, not only to organize a post-hoc narrative.
Editorial extensions
If this is right
- Teams can defend high value-if-successful on its own while honestly rating scientific certainty Low, and fund a small research stage instead of a full build.
- Causal-prediction products (what-if interventions) will systematically score lower on scientific certainty than pure predictive products unless identifying variation already exists.
- GenAI candidates get no special pass: they are ranked on the same three pillars, so novelty alone does not raise eROI.
- Before any ranking, firms must enumerate widely across analytical and generative AI; a short list often omits the best bets.
- Funding decisions should produce a mix of big wins, quick wins, strategic bets, and uncertainty-reducing research rather than only the single top-ranked project.
Reading between the lines
- The same gating logic on scientific certainty would likely demote many enterprise ‘AI for pricing’ and ‘AI for personalization’ proposals that lack a credible counterfactual design.
- Organizations that already run stage-gate or real-options processes could drop eROI’s three pillars in as the AI-specific scorecard without replacing their existing portfolio machinery.
- If coarse ordinal profiles really discriminate, the expensive practice of demanding a single pre-build ROI number from data-science teams is not just noisy but actively harmful, because it forces the catch-22 the paper describes.
- A useful external test would be whether independent raters, given only the pre-outcome facts in the Compass writeups, recover the same High vs Low likelihood split the authors report.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents eROI, a practitioner framework for prioritizing AI investments that decomposes expected ROI into three separately rated pillars — Value if Successful, Likelihood of Success, and Investment Required — each with sub-dimensions rated on a coarse Low-to-Very-High ordinal scale. The authors argue that separating value from likelihood breaks a catch-22 (ROI cannot be estimated until feasibility is known, and feasibility cannot be known without building), that likelihood should be gated by its weakest of four certainty sources (product, technical, data, scientific), and that coarse profiles suffice for ranking and portfolio assembly. The framework is illustrated on three real Compass cases from 2019, in which the authors participated as decision-makers: Likely-to-Sell recommendations (funded; later credited with nine-figure annual gross commission revenue), a Time-on-Market pricing tool (shelved, on scientific-certainty grounds rooted in causal-identification limits), and a generative-AI renovation-visualization tool (limited traction). The authors explicitly disclaim the cases as an independent test of predictive accuracy.
Significance. If the framework is as usable as claimed, the paper addresses a real and expensive problem: the RAND and industry figures cited in §1 indicate that AI project selection, not execution alone, is a major failure source. Specific strengths deserve explicit credit: (i) the manuscript is unusually candid about its own evidential limits — it states twice that the Compass cases are 'not an independent test of its predictive accuracy' and labels all business cases as back-of-the-envelope; (ii) the predictive-versus-causal distinction in §2.2, and its deployment in §4.1.4 (the LTS rollout natural experiment, with the identifying assumption stated openly) versus §4.2 (no identifying variation for price counterfactuals), is a genuinely instructive contribution rarely made this cleanly in practitioner-facing AI strategy writing; (iii) Table 1 supplies concrete rating anchors and §3.2 anticipates score inflation with concrete process safeguards; (iv) the positioning relative to RICE/ICE, real options, stage-gate, and R&D portfolio management (§2.5) is accurate and appropriately modest about novelty. The paper does not ship validated instruments, machine-checked results, or prospective tests, so
major comments (1)
- The abstract and §5 assert that 'coarse business-level ratings of the three components are enough to tell strong bets from weak ones.' This is an empirical claim about discriminative power, but the only support offered is the three §4 cases, all rated by authors who were 2019 decision-makers writing with outcome knowledge. The ratings show visible hindsight sensitivity: §4.1.1 anchors LTS's Direct Product Value at 'one extra sale per agent per year,' an anchor far easier to defend after the product succeeded, and §4.2.1 explicitly grants Time-on-Market 'its optimistic ceiling' so the rejection holds 'a fortiori' — a rhetorical move that presupposes the outcome. The manuscript concedes the cases are 'not an independent test of its predictive accuracy' (§1, §5), but the abstract states the sufficiency claim without that caveat. Please either temper the claim to match the evidence (e.g., 's
minor comments (7)
- [§4.1.4] §4.1.4: the natural-experiment evaluation of LTS is described qualitatively but no magnitudes are given (group sizes, win rates on surfaced vs not-yet-surfaced comparable contacts). The nine-figure incrementality claim thus rests entirely on the earnings-call attribution [9]. If the underlying comparison cannot be published, please state that explicitly rather than letting the description imply reported evidence.
- [§2.4 / §4.3] §2.4's rule that one fact may inform two pillars 'without being double-counted' is decision-analytically correct but likely to be misread as double jeopardy, and §4.3 is the case in point: the incumbent virtual-staging alternative lowers both Direct Product Value (incremental gain) and Product Certainty (adoption). One worked sentence showing the conditional structure (V measured conditional on adoption; L capturing the probability of adoption) would preempt the objection.
- [§3.2] The 'Research investments' archetype is labeled '(High value, High Uncertainty),' mixing vocabularies: 'Uncertainty' is not a rating on the Low–Very High certainty scale of §2.2/Table 1. Please map the archetype onto the scale (e.g., 'High value, Low-to-Medium likelihood where the binding uncertainty is reducible') so the portfolio archetypes are expressible in the framework's own terms.
- [§4.2.1 / §4.3.1] The empirical inputs ('roughly 30% of listings do not convert'; pricing accounts for 'perhaps three-quarters') are uncited. Since downstream dollar figures for both the Time-on-Market and renovation tools are built on them, please either cite sources or mark them explicitly as assumptions, consistent with the back-of-the-envelope disclaimer in §5.
- [References] Reference [29]: 'Ai isn't magical' should be 'AI isn't magical.' Also check the inline placement of footnote marker 7 in §4.3.1 ('virtual staging/renovation 7'), which appears mid-phrase.
- [Figure 2] The legend lists four favorability categories (Unfavorable / Less favorable / More favorable / Favorable); please verify in the final figure that each ordinal rating — including the in-between ratings such as 'Low to Medium' and 'Medium to High' used in the text — maps to a distinct, labeled color, since the text introduces profile ratings (e.g., 'Medium to High' for Time-on-Market value) whose cell representation is ambiguous.
- [§3.2] §3.2 recommends a cross-functional panel rating each pillar independently, then reconciling. A sentence on what the authors observed in practice — how often initial ratings diverged and whether reconciliation moved rankings — would strengthen this guidance at no evidentiary cost.
Circularity Check
No derivation-by-construction circularity; eROI is an open decision framework illustrated by author-involved retrospective cases, not a fitted or self-defining prediction chain.
full rationale
This is a managerial framework paper, not a first-principles or predictive derivation. The core relation eROI ∝ (Value if Successful × Likelihood of Success) / Investment Required is a transparent decision decomposition with stated ordinal anchors and holistic/gated aggregation rules; none of the three pillars is defined in terms of realized post-hoc ROI, so the Compass outcomes are not equal to the scores by construction. The paper openly relates eROI to prior portfolio, stage-gate, RICE/ICE, and real-options machinery rather than smuggling an ansatz or importing a uniqueness theorem from the authors. Self-citations (Compass earnings transcript, engineering blogs, Provost & Fawcett) supply factual case background and general data-science context; they are not load-bearing uniqueness or forbidden-alternative lemmas. The only mild self-reference is that the authors were 2019 decision-makers, report having used the framework then, and now rate the same three bets with outcome knowledge while disclaiming an independent predictive test—an evidential/validation limit the paper itself states, not a step where a claimed prediction reduces to its fitted input. No Eq. X = Eq. Y reduction, no parameter fit renamed as forecast, no renamed empirical law presented as forced theory. Score 1 only for that acknowledged author-involved illustration loop; central content remains independent.
Assumptions & free parameters
free parameters (3)
- Ordinal rating scale granularity (Low/Medium/High/Very High) =
4-level ordinal scale per pillar/sub-dimension
- LTS success anchor (one extra sale per agent per year) =
1 listing/agent/year; ~$25,000 commission; ~20,000 agents
- Time-on-Market and renovation lift fractions =
Various scenario fractions yielding ~$90M (ToM) and ~$50M (renovation) order-of-magnitude upsides
assumptions (5)
- domain assumption Expected attractiveness of an AI bet is monotone in Value if Successful and Likelihood of Success and antitone in Investment Required (eROI ∝ V × L / I).
- ad hoc to paper Likelihood of success is gated by the weakest of product, technical, data, and scientific certainty—any single Low caps overall likelihood.
- ad hoc to paper Coarse ordinal business ratings carry enough information to separate strong from weak AI bets without precise ROI.
- domain assumption Causal effects of asking price on time-on-market are not identified from standard observational listing data without experiments, instruments, or quasi-experimental variation.
- domain assumption Before ranking, wide enumeration of business-driven AI candidates is required; funding should target a portfolio of archetypes rather than only the single top score.
invented entities (1)
-
eROI framework (three pillars with specified sub-dimensions and combination rules)
Cite this review
Pith. "Pith review of AI Strategy: How to Choose What AI Product to Implement." pith.science (2026). https://pith.science/paper/OUYVYJY4
@misc{pith2026260723733,
author = {Pith},
title = {Pith review of: AI Strategy: How to Choose What AI Product to Implement},
year = {2026},
howpublished = {\url{https://pith.science/paper/OUYVYJY4}},
note = {Machine review of arXiv:2607.23733}
}
read the original abstract
Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommendations) flagged sales outreach opportunities and went on to account for nine figures in annual gross commission revenue. Another championed AI product (a Time-on-Market pricing tool) was rightly shelved. A simple ROI estimate could not distinguish the two. We present expected ROI (eROI), a framework that decomposes each bet into three components and rates them separately: Value if Successful, Likelihood of Success, and Investment Required. Each maps to a question executives can answer before building: How valuable would it be if it worked? How likely is it to work? And what would it cost to implement? Separating the three breaks a common catch-22: teams cannot estimate ROI until they know whether a project will work, yet cannot know whether it will work without building it. Judging Value if Successful on its own dissolves the loop, letting a team argue that a product would be valuable if it worked while it weighs how likely that is. The framework also asks, before ranking anything, whether there are enough good ideas on the table. After ranking, it guides assembling a portfolio of bets rather than funding only the single top-ranked project. We illustrate eROI on Compass's candidate AI products. Precise ROI estimates are hard to make given the inherent uncertainty of AI projects. Coarse business-level ratings of the three components are enough to tell strong bets from weak ones.
Figures
Reference graph
Works this paper leans on
-
[1]
De Bruhl, and Sydne J
James Ryseff, Brandon F. De Bruhl, and Sydne J. Newberry. The root causes of failure for artificial intelligence projects and how they can succeed: Avoiding the anti-patterns of AI. Technical Report RR-A2680-1, RAND Corporation, 2024. URLhttps://www.rand.org/pubs/ research_reports/RRA2680-1.html
2024
-
[2]
The GenAI divide: State of AI in business 2025
Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari. The GenAI divide: State of AI in business 2025. Technical report, MIT Project NANDA, July 2025. URLhttps: //nanda.media.mit.edu/ai_report_2025.pdf
2025
-
[3]
Generative AI shows rapid growth but yields mixed results
S&P Global Market Intelligence. Generative AI shows rapid growth but yields mixed results. Voice of the Enterprise: AI & Machine Learning, 2025. URL https://www.spglobal.com/market-intelligence/en/news-insights/research/2025/ 10/generative-ai-shows-rapid-growth-but-yields-mixed-results
2025
-
[4]
Now decides next: The state of generative AI in the enterprise
Deloitte. Now decides next: The state of generative AI in the enterprise. Deloitte Global Survey, 2025. URLhttps://www.deloitte.com/az/en/issues/generative-ai/ state-of-generative-ai-in-enterprise.html. Fourth-quarter wave, published January 2025. 24
2025
-
[5]
Chan Kim, Renée Mauborgne, and Mi Ji
W. Chan Kim, Renée Mauborgne, and Mi Ji. Make sure your AI strategy actually cre- ates value.Harvard Business Review, September 2025. URLhttps://hbr.org/2025/09/ make-sure-your-ai-strategy-actually-creates-value
2025
-
[6]
Politzer, and Thomas H
Melissa Valentine, Daniel J. Politzer, and Thomas H. Davenport. How to make enterprise Gen AI work.Harvard Business Review, September 2025. URLhttps://hbr.org/2025/09/ how-to-make-enterprise-gen-ai-work
2025
-
[7]
What leaders should know about measuring AI project value.MIT Sloan Management Review, February 2024
Eric Siegel. What leaders should know about measuring AI project value.MIT Sloan Management Review, February 2024. URLhttps://sloanreview.mit.edu/article/ what-leaders-should-know-about-measuring-ai-project-value/
2024
-
[8]
Gaining real business benefits from GenAI: An MIT SMR executive guide
MIT Sloan Management Review. Gaining real business benefits from GenAI: An MIT SMR executive guide. MIT Sloan Management Review Execu- tive Guide, January 2025. URLhttps://sloanreview.mit.edu/article/ gaining-real-business-benefits-from-genai-an-mit-smr-executive-guide/
2025
Show all 34 references
-
[9]
Compass Inc
Compass Inc. Compass Inc. Q4 2021 earnings call transcript. Seeking Alpha, 2022. URLhttps://seekingalpha.com/article/ 4487666-compass-inc-comp-ceo-robert-reffkin-on-q4-2021-results-earnings-call-transcript. Open-access transcript:https://www.fool.com/earnings/call-transcripts/...
2021
-
[10]
Where’s the value in AI? BCG, 2024
Boston Consulting Group. Where’s the value in AI? BCG, 2024. URLhttps://www.bcg.com/ publications/2024/wheres-value-in-ai
2024
-
[11]
Gene M. Amdahl. Validity of the single processor approach to achieving large scale com- puting capabilities. InProceedings of the April 18–20, 1967, Spring Joint Computer Conference (AFIPS), pages 483–485, 1967. doi: 10.1145/1465482.1465560
1967
-
[12]
O’Reilly Media, 2013
Foster Provost and Tom Fawcett.Data Science for Business: What You Need to Know about Data Mining and Data-Analytic Thinking. O’Reilly Media, 2013. ISBN 978-1449361327
2013
-
[13]
Paul W. Holland. Statistics and causal inference.Journal of the American Statistical Association, 81(396):945–960, 1986. doi: 10.1080/01621459.1986.10478354
1986
-
[14]
Imbens and Donald B
Guido W. Imbens and Donald B. Rubin.Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, New York, 2015. ISBN 978-0521885881. doi: 10.1017/CBO9781139025751
2015 doi
-
[15]
Causal decision making and causal effect esti- mation are not the same
Carlos Fernández-Loría and Foster Provost. Causal decision making and causal effect esti- mation are not the same. . . and why it matters.INFORMS Journal on Data Science, 1(1):4–16,
-
[16]
Joseph A. DiMasi. Risks in new drug development: Approval success rates for investigational drugs.Clinical Pharmacology & Therapeutics, 69(5):297–307, 2001. doi: 10.1067/mcp.2001. 115446
2001 doi
-
[17]
DiMasi, Ronald W
Joseph A. DiMasi, Ronald W. Hansen, and Henry G. Grabowski. The price of innovation: New estimates of drug development costs.Journal of Health Economics, 22(2):151–185, 2003. doi: 10.1016/S0167-6296(02)00126-1
2003 doi
-
[18]
RICE: Simple prioritization for product man- agers
Sean McBride. RICE: Simple prioritization for product man- agers. Intercom Blog, 2018. URLhttps://www.intercom.com/blog/ rice-simple-prioritization-for-product-managers/
2018
-
[19]
Currency, New York, 2017
Sean Ellis and Morgan Brown.Hacking Growth: How Today’s Fastest-Growing Companies Drive Breakout Success. Currency, New York, 2017. ISBN 978-0451497215
2017
-
[20]
Dixit and Robert S
Avinash K. Dixit and Robert S. Pindyck.Investment under Uncertainty. Princeton University Press, Princeton, NJ, 1994. ISBN 978-0691034102
1994
-
[21]
MIT Press, Cambridge, MA, 1996
Lenos Trigeorgis.Real Options: Managerial Flexibility and Strategy in Resource Allocation. MIT Press, Cambridge, MA, 1996. ISBN 978-0262201025
1996
-
[22]
Robert G. Cooper. Stage-gate systems: A new tool for managing new products.Business Horizons, 33(3):44–54, 1990. doi: 10.1016/0007-6813(90)90040-I
1990 doi
-
[23]
Cooper, Scott J
Robert G. Cooper, Scott J. Edgett, and Elko J. Kleinschmidt.Portfolio Management for New Products. Perseus Publishing, Cambridge, MA, 2nd edition, 2001. ISBN 978-0738205144
2001
-
[24]
Ronald A. Howard. Decision analysis: Practice and promise.Management Science, 34(6): 679–695, 1988. doi: 10.1287/mnsc.34.6.679
1988 doi
-
[25]
The economic potential of generative AI: The next productivity frontier
McKinsey Global Institute. The economic potential of generative AI: The next productivity frontier. McKinsey & Company, 2023. URL https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/ the-economic-potential-of-generative-ai-the-next-productivity-frontier
2023
-
[26]
Generate value from GenAI with “small t” trans- formations.MIT Sloan Management Review, January 2025
Melissa Webster and George Westerman. Generate value from GenAI with “small t” trans- formations.MIT Sloan Management Review, January 2025. URLhttps://sloanreview.mit. edu/article/generate-value-from-gen-ai-with-small-t-transformations/
2025
-
[27]
A practical guide to gaining value from LLMs.MIT Sloan Management Review, 66(2), 2024
Rama Ramakrishnan. A practical guide to gaining value from LLMs.MIT Sloan Management Review, 66(2), 2024. URLhttps://sloanreview.mit.edu/article/ a-practical-guide-to-gaining-value-from-llms/
2024
-
[28]
Compass, Inc
Compass, Inc. Compass, Inc. reports record fourth quarter and full-year 2025 re- sults. Press release, 2026. URLhttps://www.prnewswire.com/news-releases/ 26 compass-inc-reports-record-fourth-quarter-and-full-year-2025-results-302698918. html
2025
-
[29]
Ai isn’t magical and won’t help you reopen your busi- ness.The Wall Street Journal, 2020
Christopher Mims. Ai isn’t magical and won’t help you reopen your busi- ness.The Wall Street Journal, 2020. URLhttps://www.wsj.com/articles/ ai-isnt-magical-and-wont-help-you-reopen-your-business-11590811201
2020
-
[30]
Likely to sell recommendations for real estate
Compass True North. Likely to sell recommendations for real estate. Com- pass Engineering Blog, 2020. URLhttps://medium.com/compass-true-north/ likely-to-sell-recommendations-for-real-estate-47e2f5c37f4
2020
-
[31]
2023 profile of home buyers and sellers
Jessica Lautz et al. 2023 profile of home buyers and sellers. National Associa- tion of REALTORS, 2023. URLhttps://www.nar.realtor/research-and-statistics/ research-reports/highlights-from-the-profile-of-home-buyers-and-sellers
2023
-
[32]
Machine learning in action for Compass’s likely-to-sell recommen- dations
Compass True North. Machine learning in action for Compass’s likely-to-sell recommen- dations. Compass Engineering Blog, 2020. URLhttps://medium.com/compass-true-north/ machine-learning-in-action-for-compasss-likely-to-sell-recommendations-699a6dcd5076
2020
-
[33]
Angrist and Jörn-Steffen Pischke.Mostly Harmless Econometrics: An Empiricist’s Companion
Joshua D. Angrist and Jörn-Steffen Pischke.Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press, Princeton, NJ, 2009. ISBN 978-0691120355. 27
2009
-
[2022]
URLhttps://arxiv.org/abs/2104.04103. 25
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.