{"id":"3708662b-e3ed-4425-af3e-0df81cb42bb0","arxiv_id":"2607.26952","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CreditCardQA shows LLMs err mainly on credit-card contractual conditions and comparisons, not arithmetic, with Program-of-Thought narrowing open–closed model gaps.","lead":"The paper introduces CreditCardQA, 1,800 numerical-reasoning questions built from real credit card agreements, and shows LLMs fail more on contractual rules than on arithmetic. It matters because people already use chatbots for personal finance, and mistakes concentrate in fee and penalty edge cases that hit vulnerable users hardest.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-flagged 50-example single-model audit; that remains the load-bearing soft spot but does not overturn the measured claims.","rationale":"The paper's strongest supported claim is the multi-model CoT/PoT comparison on a new real-agreement literacy set, plus a difficulty regression that coheres with the qualitative story. Table 2 and the regression coefficients (comparisons β≈−1.33, branching and monetary negative) do not depend on the 50-example audit. The reader already conditioned the verdict on over-extension of that audit and the held-out test labels; naming inconsistency (CreditCardQA/CREDITQA/FINLITQA) is cosmetic. I find no deeper load-bearing flaw (e.g., no evidence that PoT gains are artifacts, or that questions are ungrounded). Honest non-finding on any stronger attack: leave verdict CONDITIONAL with the same rationale. Concrete multi-model re-audit would tighten or bound the failure-mode claim without changing the benchmark contribution.","tokens_in":24738,"tokens_out":603,"duration_ms":14123,"concrete_test":"Independently re-label the same 50 GPT-OSS-120B errors plus 50 stratified errors each from Gemini 3.0 Pro and Llama-3.3-70B (CoT and PoT) with two annotators and report Cohen's κ and per-type rates; if numerical-error share stays ≤20% and formula/missing-condition remain dominant across models, the taxonomy generalizes; if ranks flip or κ<0.6, the §5/abstract failure claim should be narrowed to the audited model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption is correctly identified and is the main soft spot: the claim that failures arise less from arithmetic than from misapplied rules/missed conditions (and the equity-facing edge-case narrative) rests on a manual audit of 50 incorrect GPT-OSS-120B answers (Table 3, §5) plus a single-model logistic regression (§6 / Appendix I). That sample is non-mutually-exclusive, single-model, and without reported multi-annotator agreement, so prevalence ranks (70% formula substitution, 54% missing condition, 14% numerical) and the vulnerable-user interpretation are not yet shown to generalize across the 11-model suite or closed models. However, this does not undercut the paper's primary measured results: PoT gains of +0.9–6.2 points at ±5% (Table 2), the structural difficulty of comparisons/branching/monetary features in the regression, and the existence of a real-agreement literacy benchmark. Those stand on the full eval and annotations. The concern is scope of the qualitative/social claims, not internal inconsistency of the benchmark results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces CreditCardQA (also called CREDITQA/FINLITQA), a 1,800-question numerical-reasoning benchmark built from real CFPB and related credit-card agreements, with third- and first-person variants and operator-program annotations. It evaluates 11 open- and closed-source LLMs/LRMs under CoT and PoT, reporting consistent PoT gains (about +0.9 to +6.2 points at ±5% tolerance in Table 2), competitive open-weight performance (notably GPT-OSS-120B), and a narrowing of open/closed and LRM/LLM gaps under PoT. Error analysis on 50 incorrect GPT-OSS-120B answers (Table 3) attributes failures mainly to formula/rule substitution and missed conditions rather than arithmetic; a logistic regression on structural, linguistic, and input features (§6, Eq. 1, Appendix I) identifies comparisons, branching, monetary inputs, and third-person framing as difficulty drivers, with discussion of equity-relevant edge cases.","tokens_in":24985,"tokens_out":1202,"duration_ms":24399,"significance":"If the measured results hold, the paper fills a clear gap: prior financial QA benchmarks emphasize professional/market tasks, whereas this one targets consumer financial literacy grounded in real contractual text. Strengths include a carefully constructed, agreement-sourced dataset with DSL solution programs, a multi-model CoT/PoT comparison with two numeric tolerances, and a difficulty analysis that goes beyond step count. The finding that PoT helps weaker models and that comparison/conditional structure hurts accuracy is useful for both evaluation design and practical prompting. The equity-facing interpretation is suggestive rather than fully established, but the benchmark itself is a concrete, reusable contribution for high-stakes personal-finance reasoning.","major_comments":[{"comment":"§5 and Table 3: The central qualitative claim that failures arise less from arithmetic than from misapplied rules and missed conditions rests on a manual audit of 50 incorrect GPT-OSS-120B answers (non-mutually exclusive labels, no multi-annotator agreement reported). This is too thin to support abstract/conclusion-level generalization across the 11-model suite or closed models. Either expand the audit (multiple models, agreement stats) or clearly scope the claim to this focused open-weight audit and treat cross-model prevalence as a hypothesis.","section":"§5 Error Analysis (RQ2), Table 3"},{"comment":"§6–7 and the abstract: The equity narrative (errors concentrated in late-payment/small-balance cases affecting lower-income users) is only weakly tied to the regression and the 50-example audit. Agreement-level γ effects and hardship indicators are discussed, but there is no systematic stratification of error rates by card type (subprime/secured vs premium) or by penalty/small-balance question subsets with statistical tests. Either add that analysis or tone down causal/social claims to match the evidence.","section":"§6–7, Abstract"},{"comment":"Naming and construction consistency: the manuscript alternates CreditCardQA, CREDITQA, CREDITCARDQA, and FINLITQA (§2 opening vs title/abstract/tables). This is load-bearing for a benchmark paper because release, leaderboard, and citation identity depend on a single canonical name and clear dev/test protocol (800/1000). Unify naming and state exactly what is released vs held out.","section":"§2, Table 1, Dataset Release"}],"minor_comments":[{"comment":"Table 2 header typo: \"desmonstrates\" in the CoT vs PoT discussion; also fix \"formodels\" spacing in the \"Who benefits most\" paragraph.","section":"§4"},{"comment":"Appendix B is titled \"Additional Related Works on LLM Reasoning\" but contains the question-generation pipeline; renumber/retitle to avoid collision with Appendix A.","section":"Appendix B"},{"comment":"RAG evaluation (Appendix J) uses only 140 questions and one model; briefly caveat external validity in the main text when citing it, or move the negative RAG result fully to appendix without over-claiming.","section":"Appendix J / §4"},{"comment":"Eq. (1) writes the logit in a hard-to-parse stacked form; present a standard logit equation and define γ_c(i) and third-person coding more explicitly near the equation.","section":"§6.1, Eq. (1)"},{"comment":"Flesch–Kincaid: text says \"low\" scores corresponding to college level while reporting Min/Avg 9.2/11.3; clarify directionality (higher grade level = harder) for non-specialist readers.","section":"Table 1, §2.3"}],"recommendation":"minor_revision","confidential_remarks":"Fit for COLM is reasonable: solid benchmark + prompting/difficulty analysis. The main risk is over-claiming from the 50-example audit and equity story; if the authors scope those claims, this is close to accept. No integrity red flags beyond ordinary benchmark-naming messiness."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a real dataset contribution, not another market-doc QA clone. They built 1,800 operator-annotated questions from actual CFPB/Barclays card agreements, added first-person variants, and ran a clean 11-model CoT vs PoT eval. That is new relative to FinQA/ConvFinQA/DocMath/FinanceMath/BizBench, which sit in professional or market settings.\n\nWhat works: Table 2 is coherent. PoT helps every model (+0.9 to +6.2 at ±5%), helps weaker baselines more, and lets open-weight GPT-OSS-120B sit with the closed frontier. The construction pipeline is serious—expert annotators, DSL programs, held-out test, public dev split, glossary, prompts. The logistic regression (comparisons β≈−1.33, branching and monetary inputs also hurt; first-person helps) lines up with the qualitative failures and is more useful than step-count folklore. Naming wobbles (CreditCardQA/CREDITQA/FINLITQA) are annoying but minor. Math and citation pattern look fine for a COLM benchmark paper; no circular scoring.\n\nSoft spot, in proportion: the load-bearing claim that failures are “less arithmetic than misapplied rules/missed conditions,” and the vulnerable-user edge-case narrative, rest on 50 incorrect GPT-OSS-120B answers (non-exclusive labels, no agreement stats) plus one-model regression. Directionally plausible and consistent with the regression, but prevalence ranks and equity framing do not yet generalize across the suite or closed models. RAG appendix is thin and slightly worse—fine as a negative note. Held-out test labels mean you cannot fully re-score the headline numbers from the preprint alone; they do ship data/code intent.\n\nWho it’s for: financial NLP, consumer-AI reliability, and anyone tired of equating “finance reasoning” with 10-K arithmetic. Primary measured results stand. Qualitative/social claims need a wider multi-model audit before you lean on them hard.\n\nI’d send it to referees. Engage, use the set, cite the benchmark and PoT gaps; treat the 70%/54% error mix and vulnerable-user story as hypotheses until replicated.","headline":"Solid new consumer-finance literacy benchmark with clean PoT/CoT results; the contractual-error story is directionally right but over-claimed from a 50-example single-model audit.","tokens_in":25660,"tokens_out":559,"would_cite":true,"duration_ms":12158,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"On real credit-card agreements, language models fail less at arithmetic than at rules, conditions, and contractual exceptions.","keywords":["financial literacy","numerical reasoning","credit card agreements","Program-of-Thought","Chain-of-Thought","error analysis","conditional logic","LLM evaluation"],"falsifier":"Re-run the same error taxonomy and difficulty regression on a multi-model sample of incorrect answers with independent annotators; if arithmetic or retrieval errors dominate, or if comparisons and conditions no longer predict failure, the central failure story does not hold.","tokens_in":25580,"feed_emoji":"💳","tokens_out":897,"duration_ms":17055,"temperature":0.7,"pith_summary":"This paper introduces CreditCardQA, a 1,800-question benchmark built from real credit-card agreements so that models must reason about fees, interest, minimum payments, and penalties the way consumers actually encounter them. The authors evaluate open- and closed-source language and reasoning models with both step-by-step verbal reasoning and program-style reasoning. Program-style prompting improves accuracy for every model, helps weaker models most, and shrinks the gap between open and closed systems. Error audits and a difficulty regression show that wrong answers usually come from applying the wrong formula, skipping an if-condition, or misreading agreement language—not from botching the arithmetic once the right inputs are chosen. Comparisons, branching logic, and monetary constraints are the hardest structural features, and mistakes concentrate in edge cases such as late fees and small balances that hit financially vulnerable users hardest. The practical claim is that personal-finance literacy is a high-stakes reasoning setting where contractual interpretation, not calculator skill, is the bottleneck.","feed_headline":"Credit-card AI fails on rules, not arithmetic","feed_subtitle":"A 1,800-question benchmark shows program prompting helps, but conditions and fees still trip models up","key_machinery":"CreditCardQA: 1,800 grounded question–answer pairs from real credit-card agreements, with operator-program annotations, first- versus third-person variants, and a logistic-regression difficulty model over structural operators (comparisons, branches, money inputs) plus agreement-level effects.","core_discovery":"On CreditCardQA, model failures arise less from arithmetic than from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. Program-of-Thought consistently beats Chain-of-Thought (gains of roughly 0.9 to 6.2 points at a 5% error tolerance), narrows open- versus closed-source gaps, and helps weaker baseline models most; comparisons, conditional logic, and monetary constraints are the structural features that most reduce correctness.","pith_inferences":["The same rule-and-exception brittleness likely appears in other consumer contracts (mortgages, insurance, leases) that mix caps, floors, and nested ifs.","Training or test-time methods that force explicit enumeration of triggering conditions before any arithmetic may close more of the gap than larger base models alone.","Agreement-level difficulty clusters suggest issuers could rewrite disclosures for machine readability as a consumer-protection lever."],"forward_implications":["Program-style prompting should be preferred over pure verbal chain-of-thought for consumer finance Q&A.","Benchmarks and product evals should stress comparison, branching, and fee/penalty edge cases rather than only multi-step arithmetic length.","First-person consumer phrasing can raise accuracy relative to textbook-style third-person questions.","Simple retrieval over agreements can drop accuracy when chunks omit cross-referenced conditions, so full-document or structure-aware context still matters.","Errors concentrated on late fees and small balances imply higher downside risk for lower-income users of AI finance helpers."],"fun_headline_variants":["CreditCardQA: models trip on rules and conditions, not the math","PoT lifts credit-card reasoning; fees and edge cases still break models","LLM errors stem from misread contracts more than bad arithmetic","Comparisons and late fees expose gaps in financial LM reasoning","Program prompting narrows gaps but contractual terms still stump AI"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The detailed map of why models fail rests mainly on a manual audit of fifty wrong answers from one strong open model, plus a regression fit on that same model’s outputs.","fun_headline_variants_meta":{"raw":{"variants":["CreditCardQA: models trip on rules and conditions, not the math","PoT lifts credit-card reasoning; fees and edge cases still break models","LLM errors stem from misread contracts more than bad arithmetic","Comparisons and late fees expose gaps in financial LM reasoning","Program prompting narrows gaps but contractual terms still stump AI"]},"model":"grok-4.5","effort":"low","cost_usd":0.004144,"raw_usage":{"total_tokens":1223,"prompt_tokens":744,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":41444000,"prompt_tokens_details":{"text_tokens":744,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":389,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":744,"tokens_out":90,"duration_ms":7730,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T15:59:59.751081+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same error taxonomy and difficulty regression on a multi-model sample of incorrect answers with independent annotators; if arithmetic or retrieval errors dominate, or if comparisons and conditions no longer predict failure, the central failure story does not hold.","supporting_citations":[],"review_version":1}