REVIEW 3 major objections 8 minor 48 references
Professional analysts disagree too much for single-reference grading of valuation models, and current AI agents still trail juniors on judgment.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-07-31 16:10 UTC pith:3VOTKBFQ
load-bearing objection Strong measurement paper: single-golden valuation grading is empirically broken, and the mechanical–judgment gap is real; the human bar is partly confounded by vendor provenance and time budget. the 3 major comments →
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Point-tolerance grading against a single golden analyst answer is the wrong construct for end-to-end valuation: same-company professionals already disagree so much that a single-reference rule mostly measures that disagreement. Graded instead against observed analyst practice, current agents can assemble structurally sound models but still fall short of junior-level valuation judgment.
What carries the argument
GAUGE’s three-layer observed-practice envelope (method-level sensitivity bands, industry distributions, and same-company dispersion) plus 56 facets, eight validity gates, and deterministic structural checks, aggregated into a failure-aware score φ₀ that zeros non-completions.
Load-bearing premise
The “defensible” bands come from the same vendor analyst corpus used to prove professional disagreement, so a high envelope score may only mean “looks like this vendor’s practice,” not independently correct judgment.
What would settle it
Re-run the same agent and human panels with envelopes and multi-analyst bands built from an independent broker network or market outside this vendor corpus; if the mechanical–judgment gap and agent-vs-junior ordering reverse or collapse, the central claim about judgment deficit under observed practice does not hold.
If this is right
- Valuation and other non-unique professional artifacts should not be graded by proximity to one author’s point estimates.
- Agent progress on finance should be reported separately for mechanical construction and judgment, not as one end-to-end accuracy number.
- Hard structural gates and failure-aware scoring change leaderboards relative to additive rubric averages.
- Known-groups human baselines (senior > junior > student) become a required validity check for open-ended occupational benchmarks.
- Released methodology, gated data, a frozen 48-task core, and a withheld refresh pool enable longitudinal, contamination-aware leaderboards.
Where Pith is reading between the lines
- Any domain where experts legitimately disagree—legal opinions, medical plans, strategy memos—may need the same envelope-plus-gates pattern rather than single-reference rubrics.
- Closing the 26-point mechanical–judgment gap likely needs training signals tied to practice distributions, not only more spreadsheet tool use.
- If agents keep ranking companies by risk about as well as analysts agree with each other while still snapping WACC to textbook grids, the remaining deficit is assumption craft, not cross-sectional ordering.
- Vendor-network human baselines and envelopes invite external replication before the score is treated as a universal professional bar.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper first audits the single-golden-answer assumption behind existing financial-agent benchmarks: grading one analyst-built workbook against a same-company peer under standard point tolerances yields a median score of 0.33 across 108 directed pairs (65 companies), with 92.6% below 0.70 and zero same-vintage implied-price agreement within 10%. It then introduces GAUGE, a benchmark scoring agent-built valuation models against a three-layer observed-practice envelope (per-workbook sensitivity ranges, GICS p10–p90 distributions, same-company cross-analyst dispersion) rather than a point reference, combined with 56 facets (29 deterministic, 23 LLM-judged, 4 direct-rule), eight validity gates with score ceilings, and a failure-aware score φ₀ that zeros non-completions. Validation includes a 55-participant known-groups study (seniors 88.3, juniors 66.0, students 43.2), company-grouped cross-fitting, a perturbation selectivity control, and judge vote-sampling and cross-family judge audits. Across 24 agents on a 48-task core (1,011 scored cells), the best agent scores φ₀ = 53.4, above the student mean but below every senior; all agents show a mechanical–judgment pass-rate gap (fleet median 26 points). The authors release methodology, harness, run ledgers, raw judge votes, the peer-audit JSONL, gated data splits, and a withheld refresh pool.
Significance. If the results hold, this is a useful contribution on two axes. The peer-workbook audit is a direct, well-quantified, and reusable demonstration that point-tolerance grading against one analyst reference is construct-invalid for valuation — and the underlying 632 pair/tolerance records are released as JSONL, so the central negative result is independently recomputable. The benchmark itself is engineered with unusual measurement discipline: deterministic gate ablation (11/276 orderings flip), judge vote-count ablation (τ=0.944 at k=5), a cross-family judge replication showing a family-neutral uniform shift with preserved ordering (τ=0.857), a perturbation control bounding band permissiveness vs. selectivity, bootstrap rank-stability for the failure-aware ranking, verbatim failure trajectories, and full instrument provenance including disclosed detector defects and fixes. The mechanical/judgment decomposition and the WACC grid-vs-rank-ordering diagnostic (agents match analysts' cross-company risk ordering, ρ≈0.36–0.42 vs. 0.38) localize the agent deficit in a falsifiable way. The main limitation is single-vendor provenance of both the calibration corpus and the human panel, which the
major comments (3)
- [§6.3, Table 5, Fig. 5] The known-groups panel and the envelope calibration share one vendor network (§6.3: participants 'drawn from the commercial vendor network that produced the corpus'; §5.2: E-industry/E-company bands calibrated on that corpus). The paper's own numbers sharpen the concern: seniors' mechanical subscore (92.5, Table 5) essentially equals the best agent's (93%), so the entire senior–agent gap lives in judged/envelope facets — precisely the component calibrated to the vendor's own practice distribution. The App. N.2 perturbation control cannot detect a house-style offset shared by calibration corpus and panel. Request: (i) an explicit decomposition of the human–agent gap by grader family (deterministic+gates vs. judged vs. envelope), which is computable from released artifacts; (ii) framing in §6.3/abstract that φ₀ measures conformity to this observed-practice distribution, with external repli
- [§5.2, §6.5, App. N.1–N.2] Two load-bearing envelope numbers are weaker than the headline suggests. (i) Strict held-out E-method price coverage is 53.8% (CI 38.5–68.6%) — nearly half of real peer prices fall outside the full-credit band, so E-method functions mostly as a partial-credit device. (ii) The p90 implied-price tail rests on 17 undirected pairs, and the multiplicative near band (δ≈1.0) is asymmetric: it rejects 38.5% of upward but 0% of downward ±2× counterfeits (Table 20). The authors read the near band as partial credit only, which is defensible, but then the 91.2% near-coverage figure does little evidentiary work and should not be presented alongside strict coverage as validation. Please state the asymmetry and the n=17 basis in the main text (§5.2 or §6.5), not only in App. N.
- [§5.4, App. K.1, App. N.4–N.5] App. N.4 shows the leaderboard leans most on the judged facets (τ drops to 0.683 without J; top-1 preservation 0.305), yet the only human-correctness evidence for the judge is the second-hand 460-case note (App. K.1) with no annotator qualifications, sampling frame, consensus construction, blinding, or item-level labels — and its weakest slice is valuation (82.9% exact, κ=0.75), the facet family carrying the paper's central claim. The cross-family replication (App. N.5) addresses bias but both judges could share systematic errors vs. experts. The authors' downgrade to 'descriptive' is appropriate, but for a venue publication the judgment-facet correctness claim needs either documentation of this audit or a small documented expert re-audit of valuation facets (the five lowest-agreement facets flagged in N.5 are a natural target).
minor comments (8)
- [§5.2] The §5.2 prose defining the layers is garbled: 'available for 54uses same-GICS p10 to p90 distributions ... covering 8665-company multi-coverage corpus' — numbers appear fused (54%? 866 books? 65 companies). Please rewrite the paragraph around Eq. (1).
- [§6.3, App. D.4] Human–agent comparisons use φ₀ for both populations, but humans were untimed (median 200–273 min, App. D.4) while agent timeouts score zero. For the headline claim this is largely benign (the best agents complete 100%, so their φ₀ = φ), but for mid-fleet agents it is not. A sentence noting that the human–agent ordering also holds on completed-attempt φ (Table 18's Φ column vs. Table 5) would close this.
- [§5.3, Table 18] Notation collision: Φ in Eq. (5) is the gated final score, but Table 18's Φ column is a completed-cells conditional quantity that differs from Table 4's φ₀ (e.g., Grok 46.5 vs. 40.7; Doubao 2.1-pro 40.6 vs. 10.2). Rename one, and state in the Table 18 caption that rows are not rank-ordered by it.
- [§4, Fig. 2, Fig. 9] The 'flat score' of §4/Fig. 2 is never formally defined (fraction of stated criteria passed? equally weighted?). One line defining it, and noting unstated criteria are N/A, would help. Also Fig. 2 is near-duplicate of Fig. 9a; consider merging.
- [§6.3] Experience groups are 'vendor-classified' (§6.3): the known-groups result partly validates the vendor's own seniority labels. A caveat sentence is warranted alongside the existing credential-verification disclaimer.
- [App. J, §7] 22 of 25 industry overlays were generated by claude-opus-4-8, the same model as the ladder judge (App. J). The disclosure is commendable; please also note it in §5.3 or Limitations, since overlay activation affects which judged facets enter A(x).
- [§6.1, App. N.3] One generation per agent–task cell (§6.1) with generation variance measured only via a sentinel set: given the exact top-five set is preserved in only 22.4% of bootstrap replicates (App. N.3) and the judge flip rate is 2.2%, a sentence quantifying expected rank noise near ties would calibrate reader interpretation of the leaderboard.
- [Table 15, App. T] G5 (look-ahead) is described in App. T as catching time-inverted model mechanics rather than information leakage, since packs are as-of-clamped by construction. Consider renaming or re-scoping the gate description in Table 15 to match what it actually detects.
Circularity Check
Main empirical claims are not circular; partial circularity sits in the validity apparatus (same-vendor envelope + known-groups, and corpus context improving envelope-scored facets).
specific steps
-
fitted input called prediction
[§5.2 Three-Layer Defensibility Envelope; App. N.1 cross-fit]
"eE_f(x) the same band widened by the p90 cross-analyst disagreement for that assumption (Section 4). ... E-company ... sets the near-band widths in eE_f and tests the two proxy layers. ... all folds remain within the same 65-company source corpus. These results provide internal validation of sampled practice, not external replication or proof that every in-band choice is correct."
Near-band widths and the peer-audit disagreement tails that justify them are estimated from the same multi-covered corpus used to motivate and calibrate the envelope. Held-out coverage checks are company-grouped but still resample that single source distribution, so ‘envelope admits observed analyst practice’ is partly guaranteed by construction of the bands from that practice rather than an external referent.
-
other
[§6.3 Human Baseline: Known-Groups Validity; §7 Limitations]
"Participants were drawn from the commercial vendor network that produced the corpus and grouped by the vendor’s seniority classification... The corpus and human study also come from one vendor network, with unverified credentials and overrepresentation of US listings."
Known-groups ordering is offered as primary construct-validity evidence that φ tracks valuation experience. Because judgment facets are scored against an observed-practice envelope built from that same vendor’s workbooks, senior outperformance partly measures match to the house distribution that defines the bands—not an independently anchored standard of defensibility. This is a shared-provenance validity loop, not a formal self-definition of a derived equation.
-
fitted input called prediction
[§6.6 Training Signal; Appendix O]
"E-industry tables raise the valuation-judgment subscore by +4.0 φ (95% CI [+1.3, +6.7]) in a leakage-controlled 200-workbook split... In both studies, the judgment gains concentrate on envelope-scored facets—the quantities the corpus distributions directly inform."
Supplying corpus-derived industry envelope tables as context and then measuring gains on facets graded by those same envelope distributions is a statistically forced, scorer-aligned lift. The paper correctly notes concentration on envelope-scored facets; that localization is exactly the by-construction channel.
full rationale
GAUGE is a benchmark paper, not a first-principles derivation. The load-bearing peer-audit result—one analyst workbook scored against another under fixed tolerances—is an independent empirical measurement and does not reduce to its inputs by construction. Agent leaderboards use a frozen facet/gate stack on held-out tasks and are likewise non-circular. What carries a mild circularity burden is the construct-validity loop around “defensible judgment”: E-method/E-industry bands and p90 near-widths are estimated from the same multi-covered vendor corpus that motivates abandoning single-golden grading, and company-grouped cross-fits remain inside that 65-company source sample (explicitly internal validation, not external replication). The 55-person known-groups panel is drawn from the same commercial vendor network that produced the corpus, so high senior scores partly test conformity to a house practice distribution the instrument encodes. Separately, the training-signal study shows E-industry tables raise the valuation-judgment subscore on facets the corpus distributions directly grade—an expected, partly by-construction lift the paper itself localizes. These issues weaken how far φ can be read as transferable professional judgment, but they do not make the peer-disagreement statistic or the mechanical–judgment agent gap tautological. Score 3: real but partial circularity in the referent/validity chain; central comparative claims retain independent content.
Axiom & Free-Parameter Ledger
free parameters (6)
- Base point-tolerance bands (rev ±5%, WACC ±50bp, price ±10%) and 1×–4× sweep =
rev±5%; WACC±50bp; price±10%
- Near-band width = p90 cross-analyst disagreement; industry bands at p10–p90 (operating p90) =
p90 near rule; E-industry p90
- Facet score map φ(0)=0, φ(1)=60, φ(2)=100 =
0/60/100
- Eight gate ceilings κ_g (e.g. G1→40, G5→35) =
G1:40, G2:45, G3:55, G4:50, G5:35, G6:50, G7:60, G8:60
- Judge vote count k=5 majority reduction =
k=5
- Deterministic detector thresholds (BS 0.1%, cash 0.5%, formula density 90/95%, etc.) =
as in Tables 10–15 / gates.json
axioms (6)
- domain assumption Standard three-statement accounting identities and DCF/WACC (or bank DDM) mechanics are the correct structural targets for “model construction.”
- domain assumption Independently produced vendor analyst workbooks constitute a valid sample of “observed professional practice” for envelope bands.
- domain assumption Vendor seniority classes (senior/junior/student) are ordered proxies for valuation experience in the known-groups study.
- ad hoc to paper Values inside empirical analyst envelopes are more defensible than values outside, without claiming unique correctness.
- ad hoc to paper A frozen LLM ladder judge with k-vote majority is an adequate operational grader for 23 qualitative facets.
- ad hoc to paper Non-completions and invalid artifacts should score zero on a fixed task denominator (φ₀).
invented entities (3)
-
Three-layer defensibility envelope (E-method, E-industry, E-company calibration)
independent evidence
-
56-facet taxonomy with C1/C2/C3 mechanical–judgment split and 8 validity gates
independent evidence
-
Failure-aware score φ₀ / gated aggregate Φ
no independent evidence
read the original abstract
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.
Figures
Reference graph
Works this paper leans on
-
[1]
Anthropic. 2026. Claude Fable 5 & Claude Mythos 5 System Card. https://www. anthropic.com/claude-fable-5-mythos-5-system-card. Accessed 2026-07-23
2026
-
[2]
Anthropic. 2026. Claude Opus 4.8 System Card. https://www.anthropic.com/ claude-opus-4-8-system-card. Accessed 2026-07-23
2026
-
[3]
Anthropic. 2026. Claude Sonnet 5 System Card. https://www.anthropic.com/ claude-sonnet-5-system-card. Accessed 2026-07-23
2026
-
[4]
Andrew M Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, et al. 2026. Measuring what matters: Construct validity in large language model benchmarks.Advances in Neural Information Processing Systems38 (2026)
2026
-
[5]
DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Con- text Intelligence. arXiv:2606.19348 [cs.CL]
arXiv 2026
-
[6]
Efthimios G Demirakos, Norman C Strong, and Martin Walker. 2004. What valuation models do analysts use?Accounting horizons18, 4 (2004), 221–240
2004
-
[7]
Google DeepMind. 2026. Gemini 3.1 Pro Model Card. https://deepmind.google/ models/model-cards/gemini-3-1-pro/. Accessed 2026-07-23
2026
-
[8]
Google DeepMind. 2026. Gemini 3.5 Flash Model Card. https://deepmind.google/ models/model-cards/gemini-3-5-flash/. Accessed 2026-07-23
2026
-
[9]
Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Steven Wu, and Alexandra Chouldechova. 2026. Validating llm-as-a-judge systems under rating indeterminacy.Advances in Neural Information Processing Systems38 (2026), 112282–112350
2026
-
[10]
Valentin Hofmann, David Heineman, Ian Magnusson, Kyle Lo, Jesse Dodge, Maarten Sap, Pang Wei Koh, Chun Wang, Hannaneh Hajishirzi, and Noah A Smith. 2025. Fluid language model benchmarking.arXiv preprint arXiv:2509.11106 (2025)
Pith/arXiv arXiv 2025
-
[11]
Shahed Imam, Richard Barker, and Colin Clubb. 2008. The use of valuation models by UK investment analysts.European accounting review17, 3 (2008), 503–535
2008
-
[12]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974(2024)
Pith/arXiv arXiv 2024
-
[13]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real- World GitHub Issues? arXiv:2310.06770 [cs.CL] https://arxiv.org/abs/2310.06770
Pith/arXiv arXiv 2024
-
[14]
Kimi Team. 2025. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534 [cs.LG]
Pith/arXiv arXiv 2025
-
[15]
Kimi Team. 2026. Kimi K3 Tech Blog: Open Frontier Intelligence. https://www. kimi.com/blog/kimi-k3. Accessed 2026-07-23
2026
-
[16]
Michael Krumdick, Varshini Reddy, Shivani Chaudhary, William Day, Maarij Ahmed, Hayan Haqqi, Muhammad Ahsen Fahim, Hanzallah Amjad, Ahmad Orakzai, Aqsa Gul, et al. 2026. FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks.arXiv preprint arXiv:2604.05912(2026)
Pith/arXiv arXiv 2026
-
[17]
Srivatsa Kundurthy, Clara Na, Colton Moraine, Anoushka Mohta, Case Winter, George Fang, John Ling, Emma Strubell, and Zach Kirshner. 2026. BlueFin: Bench- marking LLM Agents on Financial Spreadsheets.arXiv preprint arXiv:2605.30907 (2026)
Pith/arXiv arXiv 2026
-
[18]
Elaine Lau, Markus Dücker, Ronak Chaudhary, Hui Wen Goh, Rosemary Wei, Vaibhav Kumar, Saed Qunbar, Guram Gogia, Yi Liu, Scott Millslagle, et al. 2026. BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows. InRLEval: Methods and Reinforcement Learning Environments for Evaluating AI Agents
2026
-
[19]
Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, Xiaokang Zhang, Xiaohan Zhang, Sijia Luo, Xi Wang, and Jie Tang. 2024. Spreadsheetbench: Towards challenging real world spreadsheet manipulation.Advances in Neural Information Processing Systems37 (2024), 94871–94908
2024
-
[20]
MiniMax. 2026. MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — All in One Model. https://www.minimax.io/blog/minimax-m3. Accessed 2026- 07-23
2026
-
[21]
OpenAI. 2025. gpt-oss-120b & gpt-oss-20b Model Card. arXiv:2508.10925 [cs.CL]
Pith/arXiv arXiv 2025
-
[22]
OpenAI. 2026. GPT-5.6 System Card. https://deploymentsafety.openai.com/gpt- 5-6. Accessed 2026-07-23
2026
-
[23]
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024. tinyBenchmarks: evaluating LLMs with fewer examples. arXiv preprint arXiv:2402.14992(2024)
Pith/arXiv arXiv 2024
-
[24]
Qwen Team. 2025. Qwen3-Coder: Agentic Coding in the World. https://qwenlm. github.io/blog/qwen3-coder/. Accessed 2026-07-23
2025
-
[25]
Qwen Team. 2026. Qwen3.7-Max Model Documentation, Alibaba Cloud Model Studio. https://www.alibabacloud.com/help/en/model-studio/models. Accessed 2026-07-23
2026
-
[26]
Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J Kochenderfer. 2024. Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices.Advances in Neural Information Processing Systems37 (2024), 21763–21813
2024
-
[27]
Bytedance Seed. 2026. Seed1. 8 model card: Towards generalized real-world agency.arXiv preprint arXiv:2603.20633(2026)
Pith/arXiv arXiv 2026
-
[28]
StepFun. 2026. Step 3.7 Flash: A High-Efficiency Flash Model for Real-World Agents. https://static.stepfun.com/blog/step-3.7-flash/. Accessed 2026-07-23
2026
-
[29]
Tencent Hunyuan Team. 2026. Tencent Hunyuan Officially Releases Hy3, Ad- vancing Agent Capabilities and Deeper Product Integration. https://hunyuan. tencent.com/research/100064?langVersion=zh. Accessed 2026-07-23
2026
-
[30]
A Wang, G Meinhardt, J Katz, JH Kim, PK Chaudhary, C Blagden, and E Xu. 2026. BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents.arXiv preprint arXiv:2606.03829(2026)
Pith/arXiv arXiv 2026
-
[31]
Sinuo Wang, WANG PIAOHONG, Tianrui Qin, Maojia Song, Qianben Chen, Qiexiang Wang, Gengze Zhou, Zeyu Zhang, He Zhu, Dingfeng Shi, et al. 2026. EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evo- lution in Reinforcement Learning. InForty-third International Conference on Machine Learning
2026
-
[32]
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al. 2024. Livebench: A challenging, contamination-limited llm benchmark.arXiv preprint arXiv:2406.19314(2024)
Pith/arXiv arXiv 2024
-
[33]
xAI. 2026. Introducing Grok 4.5. https://x.ai/news/grok-4-5. Release announce- ment, July 2026
2026
-
[34]
An Yang, Anfeng Li, Baosong Yang, et al . 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL]
Pith/arXiv arXiv 2025
-
[35]
Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. 2026. Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763(2026)
Pith/arXiv arXiv 2026
-
[36]
Taojie Zhu, Wentao Zhao, Rui Sun, Beidi Luan, Jiacheng Lu, Sinuo Wang, Jing Li, Daxin Jiang, Yonghong He, and Zuo Bai. 2026. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets. arXiv preprint arXiv:2605.28359(2026)
Pith/arXiv arXiv 2026
-
[37]
922 tickers, 25 GICS industry groups
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Shu Liu, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, et al. 2026. Establishing best practices in building rigorous agentic benchmarks.Advances in Neural Information Processing Systems38 (2026). KDD ’27, August 2027, San Jose, CA, USA Appendix Appendix Contents The appendix i...
2026
-
[38]
Create the ./output/ directory if it does not exist
Save the workbook to ./output/{TICKER}_model.xlsx. Create the ./output/ directory if it does not exist
-
[39]
Produce a SINGLE workbook with all sheets in it
-
[40]
11 KDD ’27, August 2027, San Jose, CA, USA Appendix 0 0.25 0.50 0.75 1 Flat single-golden score of one analyst vs
Every projected/forecasted number must be a live Excel **formula** that references inputs -- never a value computed in Python and written as a number. 11 KDD ’27, August 2027, San Jose, CA, USA Appendix 0 0.25 0.50 0.75 1 Flat single-golden score of one analyst vs. a peer 0 20 40 60 80 100 Cumulative % of 108 pairs 1× 2× 3× 4× 0.70 92.6% of pairs below 0....
2027
-
[41]
The model must produce a final **implied share price** that is clearly labeled and easy to find
-
[42]
Use real accounting conventions (GAAP-style). The Income Statement, Balance Sheet, and Cash Flow Statement must tie together -- net income flows to retained earnings and to the cash flow statement, ending cash on CF equals cash on BS, etc
-
[43]
TODO" cells, no
No placeholder text, no "TODO" cells, no "Excel Data Table feature" notes. The workbook must be fully functional when opened. Do not ask clarifying questions. Make reasonable analyst-grade assumptions. Document them in the workbook (cell comments or an Assumptions sheet) but do not block on them. Work efficiently. Spend your reasoning on the model itself,...
-
[44]
{model_xlsx} A single institutional-quality .xlsx with these sheets: Cover, Assumptions, Revenue_Build (segment drivers), Income_Statement, Balance_Sheet, Cash_Flow, Debt_Schedule, WACC, DCF (Valuation), Sensitivity, Checks. Hard requirements (machine-graded + judged): - 5 forecast years after the last actual; every forecast cell is a live FORMULA referen...
-
[45]
{memo_md} A short investment memo: thesis with 3-5 falsifiable, quantified claims; key risks mapped to model drivers; headline numbers (implied price, EPS) that MATCH the workbook
-
[46]
{assum} JSON list of every key assumption: {"name","value","source"} where source is a resolvable pointer ("Sheet!Cell", or the input row it came from). Every number you cite in the memo must appear here. Build the workbook now. When finished, reply with ONLY the path you wrote. Do not narrate. D.3 Rendered Input Pack: Excerpt The pack is dumped sheet by ...
arXiv 2027
-
[47]
A5"] = "Cash & Equivalents
spans 2022A–2029E): Goodwill & Intangibles 15,000, Other Non- Current Assets 5,000, Other Current Assets 500 — none of these values appears anywhere in the input pack — while the pack’s gen- uine FY2024 cash figure (3,127) is copied backwards into 2022A and 2023A as well. Each constant is styled blue_font, the analyst con- vention for a legitimate hardcod...
2027
-
[494]
abstain, don’t fabricate
— so the partial-coverage limitation of Section 7 is checkable, not just confessed. The analyst-vs-analyst audit (Section 4) releases its 632 per-pair, per-tolerance records (158 directed same-ticker pairs at four tolerance multipliers), so the paper’s central negative result is recomputable from JSONL. And every judged cell in the frozen run retains all ...
2027
This paper was first reviewed by grok-4.5 on July 31, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.