{"id":"99b5d1de-7b0b-47e0-9cb6-3470efdf932b","arxiv_id":"2502.02846","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Simulations show response-scale reliability plateaus around 5 to 10 categories when item error is fixed, and peaks at 4 to 7 categories when error grows as options increase.","lead":"This study uses simulations to ask how many response options a survey item should have. It finds reliability plateaus near 5 to 10 categories when item error is fixed, and peaks at 4 to 7 categories when error grows with the number of options, arguing against casually converting Likert scales to 100-point visual analog scales.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'drastic VAS decline' rests on linear error-growth slopes (Sec 2.2) that are assumed, not estimated; the optima 7/5/4 are direct consequences of those slopes, so the practical conclusion is not empirically supported.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the error-dependency slopes are arbitrary inputs rather than empirical estimates, and the reported optima are direct consequences of those slopes. My stress-test adds that the paper extends this assumed linear relationship beyond its simulated range (k=2..20) to make claims about 100-point VAS, which makes the practical conclusion even more fragile. A concrete empirical calibration of σ(k) would settle whether the central claim lands. Because the simulation is transparent and the conditional claims are coherent, the appropriate verdict remains CONDITIONAL rather than REJECT; the concern is about the strength of the practical generalization, not the internal logic of the simulation.","tokens_in":16247,"tokens_out":4583,"duration_ms":50648,"concrete_test":"Calibrate the error-growth function empirically: in a within-person EMA experiment, administer the same item content in formats with, e.g., 4, 7, 11, 21, and 101 response categories (randomized order), and fit the authors' GRM item-variable parameterization separately per format to estimate σ_k. Plug the estimated σ_k sequence into the same simulation (thresholds -2 to 2, single and three items, N=100/500/1000) and recompute the optima and the k=101 reliability. If the empirical σ(k) is not steeply increasing or if the simulated VAS reliability drop is small, the central claim fails; if it is steeply increasing and the drop remains drastic, the claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 defines the dependent-error conditions by fiat: for k=2..20, measurement error σ_j(k) is set to 0.05–0.5 (small), 0.1–1.0 (medium), or 0.2–2.0 (large), increasing linearly. These slopes are not estimated from any data, and they are the entire driver of the reported optima (7, 5, and 4). The independent-error condition in Section 3.1 shows no optimum, which confirms that the optimum is inserted by the chosen σ(k) function. The paper then extrapolates these lines beyond k=20 to condemn VAS (0–100), even though the dependent-error simulations stop at 20 categories. The conclusion that 'conversion of any Likert scale item to VAS will result in a drastic decrease in reliability' therefore depends entirely on an unmeasured, linear-in-k error-growth assumption. The authors themselves concede in Section 5 that 'these optima are highly dependent on our simulation design.' Without empirical calibration of σ(k), the practical recommendation against VAS is not supported by the simulation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper simulates Likert- and VAS-style response data using the graded response model with an item-variable parameterization, and examines how the number of response categories affects recovery of the latent trait (Spearman correlation) and the standard error of a regression coefficient. In the independent-error condition, reliability increases and then plateaus. In the dependent-error condition, measurement error is assumed to increase linearly with the number of categories at three slopes (small, medium, large), yielding optima at 7, 5, and 4 categories, respectively. The authors conclude that converting Likert items to VAS (0-100) will drastically reduce reliability and that any such conversion requires re-validation.","tokens_in":16482,"tokens_out":3124,"duration_ms":30472,"significance":"If taken as a conditional simulation study, the paper is a useful proof of concept: it demonstrates that under a monotonically increasing error-category relationship, a finite optimum exists and the location of that optimum shifts with the error-growth rate. The item-variable construction of the GRM is a clean way to separate categorization error from item-level measurement error. However, the practical conclusions are not empirically supported because the linear error-growth slopes in Section 2.2 are chosen by fiat, not estimated from data, and the reported optima are direct consequences of those slopes. The paper is also transparent about this dependence in Section 5, but the abstract and conclusion overstate the generalizability. With appropriate reframing and additional simulation detail, the paper could be a worthwhile contribution to the response-format literature.","major_comments":[{"comment":"The optima of 7, 5, and 4 categories are direct outputs of the three linear error-dependency sequences specified in Section 2.2 (σ_j from 0.05-0.5, 0.1-1.0, and 0.2-2.0 across k=2..20). These slopes are not estimated from any empirical data, and the independent-error condition in Section 3.1 shows that no optimum arises without them. The abstract's claim that 'conversion of any Likert scale item to VAS will result in a drastic decrease in reliability' is therefore not supported by the simulations alone. The authors should reframe the contribution as a conditional demonstration, explicitly state that the optima are functions of the assumed error-growth function, and ideally include a sensitivity analysis across a range of functional forms (e.g., sublinear, concave, or stepwise) to show how the conclusion depends on the shape of σ(k).","section":"§2.2, §3.2, §5"},{"comment":"The dependent-error simulations only vary the number of response categories from 2 to 20, yet the abstract and conclusion extrapolate the 'drastic decrease' to VAS scales with 0-100 response points. This extrapolation assumes that linear error growth continues unchanged from k=20 to k=100, which is an untested assumption. The authors should either simulate the full 2-100 range under their dependency rules, or explicitly restrict the practical claim to the range of categories actually simulated and discuss what additional evidence would be needed to extend it to VAS.","section":"§2.2, §3.2, Figures 5-6"},{"comment":"The number of Monte Carlo replications is never reported, and none of the figures include error bars or confidence bands. For a study whose central claim is the location of an optimum and the existence of a decline beyond it, the reader cannot distinguish a real turning point from Monte Carlo noise. The authors should report the number of replications and add uncertainty estimates (e.g., standard errors or confidence bands across replications) to Figures 1-6, or at least report the Monte Carlo standard error for the reported optima.","section":"§2.2, all figures"}],"minor_comments":[{"comment":"The text states there is a one-to-one relationship between σ_j and the GRM discrimination parameter but does not provide the analytic mapping; giving the formula would improve reproducibility and help readers interpret the error sequences.","section":"§1.1"},{"comment":"The independent-error condition sets σ_j to values from 0.1 to 1.0, but Figure 1's legend shows only 0.25, 0.50, 0.75, 1.00; clarify whether these are the full set of levels or a selected subset, and if the latter, why the subset was chosen.","section":"§2.2"},{"comment":"The text says the SE for large dependency 'rises again after approximately five response options,' but Figure 6 appears to show the increase beginning around four categories; please align the description with the plotted results.","section":"§3.2, Figure 6"},{"comment":"The manuscript references Supplementary Materials for regression-coefficient bias figures but does not state how to access them; please include a data/code availability statement.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper's central simulation design is coherent and the conditional findings are worth publishing, but the abstract and conclusion substantially overstate the empirical basis. The lack of replication count is a reproducibility issue that should be fixed before acceptance. The authors may benefit from explicitly positioning the work as a theoretical possibility proof under an assumed error-growth mechanism, and from adding a section on how one would empirically estimate σ(k) for real items."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on Sun, Schmidt, and Henry (arXiv:2502.02846). The paper does something genuinely useful: it gives a formal framework for a question that usually gets answered by folklore. The item variable construction from Lord and Novick lets them define item-level measurement error independently of the number of response categories, and the GRM-based data generation is coherent. The independent-error results—reliability plateaus after roughly 5–10 categories, with no true optimum—are a clean baseline and match the existing simulation literature. The dependent-error simulations are the novel part: with three specified linear error-growth structures, the optima land at 7, 5, and 4 categories, and reliability declines beyond each. That conditional claim is internally consistent.\n\nThe soft spot is exactly what the stress-test note flags. The three error-growth slopes in Section 2.2 are chosen by fiat: σ(k) increases linearly from 0.05–0.5, 0.1–1.0, or 0.2–2.0 over k = 2 to 20. These slopes are not estimated from data, and the optima are direct outputs of them. The authors concede this in Section 5: 'these optima are highly dependent on our simulation design.' That honesty is creditable, but it doesn't rescue the abstract, which says converting a Likert item to VAS (0–100) 'will result in a drastic decrease in reliability.' That claim extrapolates the linear error-growth assumption beyond the simulation range of 2–20 categories to 100 categories, with no empirical support. Also missing: the number of Monte Carlo replications, error bars or confidence intervals on any figure, and code or data. Those are fixable omissions, not fatal flaws.\n\nThe discussion is more balanced than the abstract. The authors engage with Haslbeck et al. (2024) and other empirical work showing VAS can outperform Likert in EMA, and they soften the conclusion to 'problematic' rather than 'invalid.' The framework could become genuinely useful if the error-dependency function were calibrated from real data—test-retest, response consistency, or within-person studies across formats.\n\nWho gets value from this? Methodologists working on EMA and scale design, and researchers deciding between Likert and VAS. It deserves a serious referee, but the revision needs to (1) estimate or justify the error-growth function empirically, (2) run the dependent-error simulations across the full 2–100 range or explicitly limit conclusions to the range studied, and (3) report replications and uncertainty. I'd accept it for review, with my own verdict conditional until those pieces land.","headline":"A useful formal framework for a classic scale-design question, but the headline VAS warning rests on assumed error-growth slopes not estimated from data.","tokens_in":16984,"tokens_out":3225,"would_cite":false,"duration_ms":28517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P15"],"pacs":[],"model":"deepseek-v4-flash","headline":"If measurement error rises with the number of response options, the optimal Likert scale has 4-7 categories, and converting to a 100-point visual analog scale can sharply reduce reliability.","keywords":["Likert scale","Visual Analog Scale","response categories","measurement error","reliability","Graded Response Model","item variable construction","simulation"],"falsifier":"Estimate item-level measurement error directly by administering the same item text with 2, 5, 7, 11, and 101 response options to the same respondents, then check whether item discrimination or test-retest reliability declines as categories increase; if reliability stays flat or keeps improving through 100 options, the predicted reliability collapse and the VAS caution are falsified for that construct.","tokens_in":16024,"feed_emoji":"📉","tokens_out":6378,"duration_ms":50505,"temperature":0.7,"pith_summary":"This paper asks whether there is an optimal number of response categories for survey items, and where it falls. Using simulations built on an item variable construction of the graded response model, it finds that if measurement error is independent of category count, reliability rises and then plateaus with no true optimum. Under the more realistic assumption that measurement error grows as more fine-grained options are added, reliability peaks at 7, 5, or 4 categories for small, medium, or large error-growth rates, then declines. The authors conclude that converting a validated Likert item to a 100-point visual analog scale can sharply reduce reliability, and that any response-format change requires fresh validation.","feed_headline":"Study: 4 to 7 response options beat 100-point scales","feed_subtitle":"Converting Likert items to 100-point sliders can sharply cut reliability, a simulation study finds.","key_machinery":"The item variable construction of the Graded Response Model (GRM) is the mechanism that carries the argument. Each item has a continuous latent response variable $\\gamma_{ij} \\sim N(\\theta_i, \\sigma_j)$, where $\\sigma_j$ is the item's measurement error on the latent scale; observed category choices are obtained by thresholding $\\gamma_{ij}$. This reparameterization of Samejima's GRM lets the authors define measurement error independently of the response format, then impose different linear dependencies between $\\sigma_j$ and the number of categories and trace how the optimal category count shifts.","core_discovery":"The paper's central claim is that the optimal number of response options is governed by how item measurement error depends on category count. Using the item variable construction of the graded response model, the authors simulate responses across 2 to 100 categories. When error is independent of the number of categories, recovery of the true score improves and then plateaus after roughly 5 to 10 categories, so there is no single optimum. When error increases linearly with the number of categories, a clear optimum emerges at 7, 5, and 4 response options for small, medium, and large dependency structures, with reliability declining past that point; the standard error of a regression coefficient follows the same pattern. The paper therefore argues that a VAS, considered as a 101-point Likert scale, will typically have lower reliability than a shorter Likert version of the same item, and that a format change should trigger re-validation.","pith_inferences":["The specific optima of 7, 5, and 4 are direct consequences of the chosen linear error-growth slopes; real-world error could grow nonlinearly, so the robust qualitative prediction is that an optimum exists well below VAS-like counts.","A direct empirical test would estimate item discrimination or test-retest reliability for the same item text under 2, 5, 7, 11, and 101 response options, which would map the true dependency structure and identify construct-specific optima.","For ecological momentary assessment, where single-item measures dominate, the reliability penalty of VAS conversion may be largest, but momentary-state constructs might exhibit flatter error growth than trait measures, making construct type a moderator worth testing."],"forward_implications":["If measurement error grows with category count, reliability peaks at 7, 5, or 4 response options depending on the growth rate, then declines beyond that point.","Converting an existing Likert item to a 100-point visual analog scale will decrease reliability when such error growth is present.","When measurement error is independent of category count, there is no true optimum; reliability just increases and then plateaus.","Changing the response format of a validated measure requires re-validation because the measurement error of the scale is likely to change.","Adding more items (three instead of one) improves true-score recovery but does not move the optimal number of response categories."],"supporting_citations":[{"why":"Supplies the item variable construction that defines item measurement error independently of response format.","marker":"Lord and Novick (1968)"},{"why":"Provides the graded response model that the ogive item characteristic curves are based on.","marker":"Samejima (1969)"},{"why":"Empirical finding that 7-point scales show the best psychometric properties with declines beyond 10 categories; a key comparison for the simulation results.","marker":"Preston and Colman (2000)"},{"why":"Simulation study finding optimal range of 4 to 7 categories; a key precedent for the reported optima.","marker":"Lozano, García-Cueto, and Muñiz (2008)"},{"why":"Empirical demonstration that a VAS version of the Big Five Inventory was inferior to a 6-option Likert version; supports the paper's anti-VAS conclusion.","marker":"Simms, Zelazny, Williams, and Bernstein (2019)"},{"why":"Empirical counterpoint showing VAS outperformed 7-point Likert in EMA; the paper discusses this as a challenge to interpret.","marker":"Haslbeck et al. (2024)"},{"why":"Classic result on limited information capacity that motivates the claim that respondents struggle to distinguish many response options.","marker":"Miller (1956)"},{"why":"Shows how the number of response categories attenuates correlations, used to explain why VAS may appear advantageous in correlation-based analyses.","marker":"Olsson, Drasgow, and Dorans (1982)"}],"fun_headline_variants":["Optimal response count depends on measurement error","Likert scales often beat 100-point sliders on reliability","More response options don't always improve reliability","Converting to visual analog scales may cut reliability","Best number of response options varies by error structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that an item's measurement error grows linearly with the number of response categories at the specific rates used in the simulation; if the true relationship is flat, nonlinear, or varies by construct, the predicted optima and the warning against VAS conversion do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Optimal response count depends on measurement error","Likert scales often beat 100-point sliders on reliability","More response options don't always improve reliability","Converting to visual analog scales may cut reliability","Best number of response options varies by error structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1341,"prompt_tokens":970,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":300}},"tokens_in":586,"tokens_out":371,"duration_ms":3803,"temperature":1.0,"reasoning_tokens":300,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T10:52:55.050776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate item-level measurement error directly by administering the same item text with 2, 5, 7, 11, and 101 response options to the same respondents, then check whether item discrimination or test-retest reliability declines as categories increase; if reliability stays flat or keeps improving through 100 options, the predicted reliability collapse and the VAS caution are falsified for that construct.","supporting_citations":[{"cited_title":"\\ Novick, M R","cited_arxiv_id":null,"evidence_quote":"Supplies the item variable construction that defines item measurement error independently of response format."},{"cited_title":"APACrefauthors \\ 1969","cited_arxiv_id":null,"evidence_quote":"Provides the graded response model that the ogive item characteristic curves are based on."},{"cited_title":", García-Cueto, E","cited_arxiv_id":null,"evidence_quote":"Simulation study finding optimal range of 4 to 7 categories; a key precedent for the reported optima."},{"cited_title":", Zelazny, K","cited_arxiv_id":null,"evidence_quote":"Empirical demonstration that a VAS version of the Big Five Inventory was inferior to a 6-option Likert version; supports the paper's anti-VAS conclusion."},{"cited_title":", Martínez, A J","cited_arxiv_id":null,"evidence_quote":"Empirical counterpoint showing VAS outperformed 7-point Likert in EMA; the paper discusses this as a challenge to interpret."}],"review_version":1}