{"id":"75c134df-fc0c-4423-8a8d-84516bc85c27","arxiv_id":"1908.04110","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A unified framework showing when maximum approximated likelihood estimators are consistent and asymptotically normal, with explicit rates for required integration points for simulation, QMC, Gaussian quadrature, and sparse grids.","lead":"This paper proves general conditions under which maximum likelihood estimation with approximated likelihoods remains consistent and asymptotically normal. It shows that accurate quadrature methods need far fewer integration points than simulation, for instance about log(n) instead of n.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorems 7–8 require uniform quadrature error over the whole covariate space Z; in the paper's own mixed-logit example with Z=R this sup-norm condition fails for fixed r, so the central result does not cover the headline example without an unstated compact-support assumption.","rationale":"I read the paper as establishing sufficient conditions under which MAL estimators inherit MLE asymptotics, then checking those conditions for mixed logit, Butler–Moffitt, and simulation methods. The formal statements are conditional, so the main risk is not an invalid proof step but an unstated limitation: the key uniform error E(r) is taken over the data space, and the paper's examples do not satisfy that condition when covariates are unbounded. This matches the reader's weakest assumption. My independent check of Example Ia confirms the issue concretely: finite Gauss-Hermite rules have a nonvanishing asymptotic error as the covariate tends to infinity, so sup_z convergence cannot hold; gradient terms can even make E(r) infinite. This does not destroy the abstract framework—for compact Z the conditions can hold—but it means the headline 'specific examples' claim is overbroad. I therefore recommend keeping the CONDITIONAL verdict rather than accepting or rejecting outright; the authors should either require compact Z, replace sup_z by an assumption tied to the sampling distribution (e.g., uniform convergence over compacta plus control of tails), or restrict the examples accordingly.","tokens_in":23149,"tokens_out":9983,"duration_ms":114696,"concrete_test":"Analytic check for Example Ia: fix any r-point Gauss-Hermite rule (w_j,v_j), fix θ=(0,1), and evaluate L_∞(z):= |\\tilde f_r(z,θ)−f(z,θ)|. Compute lim_{z→∞} L_∞(z)=|Σ_j w_j 1(v_j>0)−1/2|: the true limit is 1 and the quadrature limit is the mass of nodes with v_j>0. If this limit is nonzero, sup_z |\\tilde f_r−f| does not vanish, so Theorem 7(iii) fails. Also compute ∂\\tilde f_r/∂θ1 = Σ_j w_j z σ'(z v_j) at θ=(0,1); for rules with v_j=0 this grows like z, and more generally the sup over θ near (0,1) is unbounded, so E(r)=∞. Running this calculation with standard Gauss-Hermite nodes settles that an added compact-support condition on Z, or a different empirical formulation of E(r), is necessary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equation (3.7) defines E(r) as a supremum over z∈Z and θ∈Θ, and Theorem 8(ii) requires sqrt(n)E(R(n))→0. For Example Ia (Section 5, Eq. (5.2)), f(z,θ)=∫ σ(z(θ2 v+θ1)) φ(v) dv and \\tilde f_r is Gauss-Hermite quadrature. For fixed r, as z→∞, f(z,θ)→1 while \\tilde f_r(z,θ)→Σ_{v_j>-θ1/θ2} w_j = W_r(θ), so sup_z |\\tilde f_r−f| ≥ |1−W_r(θ)|, generically a positive constant for every finite r. Hence Theorem 7(iii) (lim_{r→∞} sup_{z,θ}|\\tilde f_r−f|=0) fails for unbounded Z. The gradient term is worse: ∂\\tilde f_r/∂θ1 = Σ w_j z σ'(z(θ2 v_j+θ1)). Unless no quadrature node ever equals −θ1/θ2, which cannot be guaranteed uniformly over compact Θ, the supremum over z and θ is unbounded, so E(r)=∞. The paper never imposes compactness of Z; Section 3.1 only says Z⊂R^d. Lemma 20 bounds derivatives with respect to θ for fixed z, not uniformly over z. Thus the central theorems, as stated, do not apply to the paper's motivating mixed-logit example with unbounded covariates. This is a missing scope condition, not a contradiction inside the theorems.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops asymptotic theory for maximum approximated likelihood (MAL) estimators, in which each log-likelihood contribution is replaced by a numerical quadrature approximation with an accuracy level r that grows with the sample size n. It states general consistency and asymptotic normality results for MAL estimators (Theorems 7 and 8), derives sufficient growth rates for r under algebraic and exponential quadrature error bounds (Theorem 9), and applies the framework to Monte Carlo, quasi-Monte Carlo, Gauss-Hermite quadrature, and sparse grid integration. The theoretical results are illustrated with mixed logit and Butler-Moffitt examples and a simulation study comparing approximation methods.","tokens_in":23563,"tokens_out":8226,"duration_ms":90642,"significance":"If the stated conditions are met, the paper gives a useful unifying framework: maximum simulated likelihood emerges as a special case, and the link-function results (e.g., logarithmic growth of r for exponentially convergent quadrature) are practically relevant. The algebraic manipulations in Theorem 9 are correct, and the proofs of the general theorems follow standard extremum-estimator arguments. The paper also provides a helpful simulation comparison of approximation methods. However, the scope of the central theorems is narrower than the examples suggest: the uniform sup-norm error condition used throughout is not satisfied for the paper's own mixed logit examples when covariates are unbounded, which is the standard empirical setting. This gap is load-bearing and requires a substantial clarification or modification of the assumptions.","major_comments":[{"comment":"For the mixed logit example with unbounded covariates, the uniform approximation error E(r)=sup_{z∈Z,θ∈Θ}(|\\tilde f_r(z,θ)-f(z,θ)|+||∇_θ \\tilde f_r(z,θ)-∇_θ f(z,θ)||) does not vanish as r→∞. With f(z,θ)=∫ σ(z(θ2 v+θ1)) φ(v) dv and Gauss-Hermite quadrature, for fixed r and z→∞, \\tilde f_r(z,θ) tends to W_r(θ)=Σ_{j: θ2 v_j+θ1>0} w_j while f(z,θ)→1, so sup_z |\\tilde f_r-f| is bounded below by a positive constant for generic θ. The gradient term is worse: ∂\\tilde f_r/∂θ1=Σ_j w_j z σ'(z(θ2 v_j+θ1)) has no finite supremum over z∈R when θ2 v_j+θ1=0 for some node, and over a compact parameter set containing points arbitrarily close to such a hyperplane the supremum is unbounded. Thus Theorem 7(iii) and Theorem 8(ii) fail for the paper's headline example unless Z is assumed bounded or the error norm is changed to a data-dependent or weighted norm. This is a missing scope condition, not a contradiction inside the theorems, but it must be addressed for the examples to be covered by the theory.","section":"Section 5, Example Ia; Theorem 7(iii) and Eq. (3.7)"},{"comment":"The proof of Theorem 8 defines C1(f) and C2(f) using sup_{z∈Z,θ∈Θ}||∇_θ f(z,θ)|| and sup_{z∈Z,θ∈Θ}||∇_{θθ} f(z,θ)||, and then bounds the approximation error of the log-likelihood gradient by C1(f)√n E(R(n)). However, the assumptions of Theorem 8 do not state that these suprema are finite. If Z is unbounded, as in the mixed logit example, the suprema are infinite, and the displayed inequality is not justified. The proof therefore implicitly assumes uniform boundedness of f and its θ-derivatives up to order two over Z×Θ; this assumption needs to be stated explicitly and verified in the applications.","section":"Proof of Theorem 8, constants C1(f) and C2(f)"},{"comment":"The verification of the Gauss-Hermite error bounds for the logit examples is incomplete. Lemma 20 provides bounds of the form |D^{(θ)}_α D^{(v)}_β f(θ,v)| ≤ c(α,β) θ^β e^{v^T v/2} ∏ sqrt(1+v_i^2), so the constant depends on the parameter θ through the factor θ^β. When applied to φ(v,z_i,θ)=σ(z_i(θ2 v+θ1)), the relevant parameter vector includes z_i, so the constant c(k,α) is actually a function of z_i and is not uniform over an unbounded covariate space Z. Consequently, the paper does not establish the uniform-in-z bound (4.10) or (4.13) that is needed to obtain a finite, r-independent constant in the quadrature error estimate (4.11) for the mixed logit example.","section":"Section 5 and Lemma 20"}],"minor_comments":[{"comment":"The conclusion of Theorem 2 states 'plim_{n→∞} \\hatθ_M = θ0'; this should refer to the approximated estimator \\hatθ_AM, not the infeasible M-estimator \\hatθ_M.","section":"Theorem 2"},{"comment":"The sentence beginning 'accuracy parameter r to the number of observations n we introduce a function R' lacks a subject or transition; it should read something like 'To link the accuracy parameter r to the number of observations n, we introduce a function R: N→N.'","section":"Section 2, after Theorem 2"},{"comment":"The condition numbering in Corollary 11 jumps from (ii) to (iv); condition (iii) is missing.","section":"Corollary 11"},{"comment":"The vertical axis of Figure 6.3 is labeled 'n × max abs Err', but the text and the caption refer to √n E(R(n)); the label should be √n × max abs Err to match the quantity being plotted.","section":"Section 6, Figure 6.3"},{"comment":"The word 'independed' should be 'independent'.","section":"Section 1, paragraph 2"}],"recommendation":"major_revision","confidential_remarks":"The paper's theorems are standard in structure and appear correct under stronger assumptions, but the scope gap between the abstract conditions and the headline mixed logit examples is substantial. A revision should either impose compact covariate support in the examples, prove a variant with a weighted or data-dependent error norm, or clearly state that the uniform sup-norm theory does not apply to unbounded-covariate models. With that fixed, the paper would be a useful contribution to the econometrics literature on likelihood approximation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThis is a useful paper and the core is right, but the headline examples outrun the theorems. What's genuinely new: a general MAL framework where consistency and asymptotic normality are expressed through a single worst-case approximation error E(r) that includes the function and its gradient. That is a clean abstraction. It nests MSL as a special case and gives explicit link functions R(n) for Monte Carlo, QMC, Gaussian quadrature, and sparse grids. The algebra in Theorem 9 is correct, and the simulation study is honest and well matched to the theory.\n\nThe soft spot is real and located in the scope conditions. Theorem 7 requires f ≥ δ > 0 on all of Z×Θ, and Theorem 8 requires E(r) = sup_{z∈Z, θ∈Θ} |\\tilde f_r − f| + ||∇\\tilde f_r − ∇f|| to go to zero uniformly. Those are strong. In the paper's own mixed-logit example (Example Ia), Z = R. For fixed r, as z → ∞ the true f tends to 1 while the Gauss–Hermite approximation converges to W_r(θ), the total weight of nodes above a threshold, which is generically not 1. So sup_z |\\tilde f_r − f| ≥ |1 − W_r(θ)| > 0 for every finite r. The gradient term is worse: its z-dependence makes the supremum infinite. The verification in Section 5 only checks smoothness of the integrand in v, not uniformity over z. So as stated, the theorems do not cover the motivating example. The proof of Theorem 8 also silently uses sup_{z,θ} ||∇f|| and sup_{z,θ} ||∇θθ f|| to define C1 and C2; those need to be finite, which again fails for unbounded covariates.\n\nThis is not a contradiction inside the theorems—they are true under the stated conditions—it is a missing scope condition. The fix is straightforward: either assume Z is bounded (or covariates have compact support), or replace the sup-norm with a weighted sup-norm or an L_p(μ) norm under the data distribution, which is what the log-likelihood average actually needs. Kristensen and Salanié (2017) is a related reference for exactly this kind of relaxation.\n\nWho is this for? Anyone working on theory for approximated likelihood estimators, and practitioners who want to know when deterministic quadrature is asymptotically safe. I would send it to a serious referee; the framework deserves to be in the literature, but it needs the scope conditions stated honestly before it can serve as blanket justification for deterministic quadrature.\n\nRecommendation: engage with it, but tell the authors to fix the sup-norm gap before publication.","headline":"A clean maximal-approximated-likelihood framework whose central theorems require uniform approximation error over the whole covariate space, a condition the paper's own mixed-logit example with unbounded covariates fails, so the examples outrun the theorems as stated.","tokens_in":23997,"tokens_out":3492,"would_cite":false,"duration_ms":38058,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62E20","65D30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that maximum approximated likelihood estimators have the same asymptotic distribution as exact maximum likelihood whenever the approximation's worst-case error, including the gradient, shrinks faster than $n^{-1/2}$.","keywords":["maximum approximated likelihood","maximum simulated likelihood","quadrature","quasi-Monte Carlo","sparse grids","asymptotic normality","mixed logit","numerical integration"],"falsifier":"Take a model with a known true parameter and approximate the likelihood by Monte Carlo with $r=n^\\beta$ for $\\beta=1/2$ and $\\beta=1$. The theorem predicts that the second estimator is asymptotically normal with variance $I^{-1}$, while the first uses a link function for which $\\sqrt{n}E(R(n))$ does not vanish; a simulation showing the opposite would refute the claimed rate threshold. A sharper check is to compute the actual worst-case error $E(R(n))$ for a Gauss-Hermite approximation of a mixed-logit likelihood with bounded covariates and verify numerically whether $\\sqrt{n}E(R(n))$ tends to zero under the prescribed $R(n)$ while the estimator's distribution matches $N(0,I^{-1})$.","tokens_in":22964,"feed_emoji":"📈","tokens_out":13843,"duration_ms":131970,"temperature":0.7,"pith_summary":"The paper asks whether replacing an intractable likelihood by a numerical approximation—a quadrature rule, quasi-Monte Carlo points, or a sparse grid—changes the statistical conclusions. Its answer is that the approximation error is asymptotically irrelevant precisely when the worst-case error, measured jointly on the likelihood and its gradient, disappears faster than $1/\\sqrt{n}$ as the number of integration points grows with the sample size. Under that condition the maximum approximated likelihood estimator is consistent and asymptotically normal with the same covariance matrix as the infeasible maximum likelihood estimator. The paper also translates the condition into growth requirements for the number of integration points, and shows that exponentially convergent rules need only logarithmically many points, in contrast to Monte Carlo simulation which needs roughly as many points as observations. This matters for applied work because many discrete-choice and random-effects models have likelihoods defined by integrals that cannot be computed in closed form.","feed_headline":"Approximate MLE can match exact MLE if errors decay fast","feed_subtitle":"The paper proves that quadrature errors shrinking faster than 1/sqrt(n) keep the estimator's distribution unchanged.","key_machinery":"The central object is the worst-case approximation error $E(r)$ defined in equation (3.7), together with the link function $R:\\mathbb{N}\\to\\mathbb{N}$ that couples the number of integration points to the sample size. The proof route is to view $\\hat\\theta_{\\mathrm{MAL}}$ as an extremum estimator with an approximated objective and transfer uniform convergence of the approximation to the log-likelihood and its first two derivatives; a technical lemma on logarithms bounds the distance between $\\log\\tilde f$ and $\\log f$ by a multiple of the distance between $\\tilde f$ and $f$, provided each is bounded away from zero. What carries the argument is that the condition $\\sqrt{n}E(R(n))\\to 0$ is the natural rate threshold: it makes the approximation error in the score negligible against the $O_p(n^{-1/2})$ size of the exact score.","core_discovery":"The central claim is that maximum approximated likelihood (MAL) estimation inherits the statistical properties of maximum likelihood once the approximation is good enough relative to the sample size. Define $E(r)=\\sup_{z\\in\\mathcal{Z},\\theta\\in\\Theta}(|\\tilde f_r(z,\\theta)-f(z,\\theta)|+\\|\\nabla_\\theta \\tilde f_r(z,\\theta)-\\nabla_\\theta f(z,\\theta)\\|)$. If $R(n)$ quadrature points are used for a sample of size $n$, and $\\sqrt{n}E(R(n))\\to 0$ in probability, then $\\sqrt{n}(\\hat\\theta_{\\mathrm{MAL}}-\\theta_0)\\xrightarrow{d} N(0,I^{-1})$, the same limit the exact maximum likelihood estimator would have. The paper proves this by treating the approximated log-likelihood as an M-estimator objective and showing that the approximation error in the gradient is controlled by $E(R(n))$; the same argument gives consistency from uniform convergence of $\\tilde f_{R(n)}$ to $f$. Rate conclusions follow: algebraic convergence $E(r)\\le cr^{-s}$ requires $R(n)\\sim n^{\\gamma/s}$ for any $\\gamma>1/2$, while exponential convergence $E(r)\\le c e^{-\\alpha r^\\beta}$ requires only $R(n)\\sim(\\log n)^{1/\\beta}$.","pith_inferences":["The sup-norm formulation suggests a testable modification: replacing the supremum over an unbounded data space by an expectation-weighted norm might extend the theorem to mixed-logit settings with unbounded covariates, where the current worst-case condition is infinite.","The framework likely transfers to other approximate M-estimators such as simulated GMM or minimum distance, since the underlying argument only uses uniform convergence of the approximated objective and its derivatives.","The rate threshold suggests an adaptive implementation: a practitioner could estimate the quadrature error and keep increasing $r$ until the error is below $c/\\sqrt{n}$, instead of committing to a fixed link function.","The finite-sample simulations imply that efficiency comparisons among approximation methods can be made before fitting the model, by checking which rule has the smaller worst-case error at the chosen $R(n)$."],"forward_implications":["For deterministic rules with exponential convergence, $R(n)$ as small as $(\\log n)^{1/\\beta}$ suffices, so the computational cost is roughly $n\\log n$ integrand evaluations instead of $n^2$ for Monte Carlo under identical sampling.","Quasi-Monte Carlo rules with error $O(r^{-1+\\varepsilon})$ require $R(n)=n^\\beta$ with $\\beta>1/2$, a square-root reduction compared with Monte Carlo.","Gaussian quadrature and sparse-grid approximations can therefore deliver estimators with the same asymptotic efficiency as exact maximum likelihood at substantially lower cost, provided the integrand has the required smoothness.","The paper's conditions also imply a practical warning: for a fixed approximation accuracy, the extra estimation error grows with the sample size, so large datasets demand more accurate quadrature than small ones."],"supporting_citations":[{"why":"Supplies the M-estimator consistency and asymptotic normality theorems that the paper extends to approximated objectives.","marker":"Newey and McFadden (1994)"},{"why":"Gives abstract conditions on approximated log-likelihoods; the paper's contribution is to make them checkable for integration rules.","marker":"Ackerberg et al. (2009)"},{"why":"Establishes the classical MSL rates that the new framework recovers as a special case.","marker":"Hajivassiliou and Ruud (1994)"},{"why":"Supplies the Gaussian quadrature convergence rates used to derive the required $R(n)$.","marker":"Davis and Rabinowitz (2007)"},{"why":"Provides the Gauss-Hermite error bound for unbounded domains needed for the logit examples.","marker":"Smith et al. (1983)"},{"why":"Supplies sparse-grid convergence rates used for multivariate quadrature.","marker":"Gerstner and Griebel (1998)"},{"why":"Provides the Hardy-Krause variation framework behind the quasi-Monte Carlo error bounds.","marker":"Niederreiter (1992)"},{"why":"Gives the sparse-grid Gauss-Hermite error bound on unbounded domains used for the multivariate mixed logit.","marker":"Zhang et al. (2013)"}],"fun_headline_variants":["Approximate MLE matches exact MLE if error decays faster than 1/√n","MAL estimators inherit MLE asymptotics when errors decay fast enough","Approximate likelihood: same limit as MLE if error goes to zero fast","Fast error decay makes approximate MLE as efficient as exact MLE","Approximate MLE: same asymptotic variance if approximation error is o(1/√n)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument hangs on the assumption that the approximation error has a finite worst case over every possible data point and every parameter value, and that this worst case shrinks uniformly; with unbounded covariates this is not available for the mixed-logit example, because derivatives of the logit grow linearly in the covariate.","fun_headline_variants_meta":{"raw":{"variants":["Approximate MLE matches exact MLE if error decays faster than 1/√n","MAL estimators inherit MLE asymptotics when errors decay fast enough","Approximate likelihood: same limit as MLE if error goes to zero fast","Fast error decay makes approximate MLE as efficient as exact MLE","Approximate MLE: same asymptotic variance if approximation error is o(1/√n)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001339,"raw_usage":{"total_tokens":5429,"prompt_tokens":918,"completion_tokens":4511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":4408}},"tokens_in":534,"tokens_out":4511,"duration_ms":35172,"temperature":1.0,"reasoning_tokens":4408,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:51:38.052870+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a model with a known true parameter and approximate the likelihood by Monte Carlo with $r=n^\\beta$ for $\\beta=1/2$ and $\\beta=1$. The theorem predicts that the second estimator is asymptotically normal with variance $I^{-1}$, while the first uses a link function for which $\\sqrt{n}E(R(n))$ does not vanish; a simulation showing the opposite would refute the claimed rate threshold. A sharper check is to compute the actual worst-case error $E(R(n))$ for a Gauss-Hermite approximation of a mixed-logit likelihood with bounded covariates and verify numerically whether $\\sqrt{n}E(R(n))$ tends to zero under the prescribed $R(n)$ while the estimator's distribution matches $N(0,I^{-1})$.","supporting_citations":[{"cited_title":"McFadden (1994) ‘Large sample estimation and hypothesis testing.’ vol","cited_arxiv_id":null,"evidence_quote":"Supplies the M-estimator consistency and asymptotic normality theorems that the paper extends to approximated objectives."},{"cited_title":"Geweke, and J","cited_arxiv_id":null,"evidence_quote":"Gives abstract conditions on approximated log-likelihoods; the paper's contribution is to make them checkable for integration rules."},{"cited_title":"Ruud (1994) ‘Classical estimation methods for LDV models using simulation.’ In ‘Handbook of Econometrics,’ vol","cited_arxiv_id":null,"evidence_quote":"Establishes the classical MSL rates that the new framework recovers as a special case."},{"cited_title":"Rabinowitz (2007)Methods of Numerical IntegrationDover Books on Mathematics Series (Dover Publications)","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian quadrature convergence rates used to derive the required $R(n)$."},{"cited_title":"Sloan, and A","cited_arxiv_id":null,"evidence_quote":"Provides the Gauss-Hermite error bound for unbounded domains needed for the logit examples."},{"cited_title":"Griebel (1998) ‘Numerical integration using sparse grids.’Numerical Algorithms 18, 209–232","cited_arxiv_id":null,"evidence_quote":"Supplies sparse-grid convergence rates used for multivariate quadrature."},{"cited_title":"(1992) Random Number Generation and Quasi-Monte Carlo Methods (SIAM, Philadelphia)","cited_arxiv_id":null,"evidence_quote":"Provides the Hardy-Krause variation framework behind the quasi-Monte Carlo error bounds."},{"cited_title":"Gunzburger, and W","cited_arxiv_id":null,"evidence_quote":"Gives the sparse-grid Gauss-Hermite error bound on unbounded domains used for the multivariate mixed logit."}],"review_version":1}