{"id":"532f51c3-4937-4c79-89d6-df1a9c2b5031","arxiv_id":"2605.24854","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ReLU neural network estimators for nonparametric regression with repeated measurements under covariate shift achieve minimax optimal rates, with a novel approximation theory featuring polynomial rather than exponential dependence on dimension.","lead":"This paper develops neural network estimators to predict outcomes in a target setting with repeated measurements but no observable responses, by borrowing from a source setting and adjusting for differences in input distributions via density ratios. A smart generalist might read it to see how existing data can be reused when new outcome collection is expensive or impossible.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption correctly isolates the single point at which the optimality claim could fail. Because the provided abstract and claim summary contain no contradictory or circular reasoning, and the paper explicitly claims to supply the required approximation result, the existing UNVERDICTED verdict does not need adjustment.","tokens_in":1759,"tokens_out":263,"duration_ms":15715,"concrete_test":"Extract the statement and proof of the novel approximation theorem (likely in the theoretical analysis section); verify that every constant multiplying the network width, depth, or weight bound is at most polynomial in the input dimension d (e.g., O(d^k) for fixed k independent of the smoothness index).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on establishing non-asymptotic bounds that achieve minimax rates via a new ReLU approximation result whose constants scale polynomially (rather than exponentially) in dimension. The paper states that the target regression function and density ratio lie in classes admitting such approximation under the given smoothness conditions, with separate handling for bounded versus finite-moment density ratios. No internal inconsistency, hidden exponential factor, or unsupported step in the rate derivation is visible from the stated claims and standard assumptions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript develops a transfer learning framework for nonparametric regression with repeated measurements under covariate shift, leveraging a source domain with observable responses. It proposes ReLU feedforward neural network estimators for the target regression function (and density ratio when unknown), handling both uniformly bounded and finite-moment density ratio cases. Non-asymptotic error bounds are established that achieve minimax optimal rates under the repeated measurements setting, supported by a new approximation theory in which network parameter constants scale polynomially (rather than exponentially) with dimension.","tokens_in":1865,"tokens_out":400,"duration_ms":20113,"significance":"If the central claims hold, the work would advance nonparametric regression under covariate shift by delivering sharper stochastic error bounds via improved ReLU approximation theory that mitigates the curse of dimensionality. The explicit handling of known/unknown density ratio and bounded/unbounded cases, together with the minimax optimality result, represents a substantive theoretical contribution; the numerical simulations and real-data example provide supporting empirical evidence.","major_comments":[{"comment":"Theoretical analysis section: the novel approximation result asserting polynomial (rather than exponential) dependence of network constants on dimension is load-bearing for both the sharper non-asymptotic stochastic error bounds and the claimed mitigation of the curse of dimensionality. The precise statement of the theorem (including the dependence on smoothness indices, network depth/width, and the precise function classes for the regression function and density ratio) must be given explicitly, together with the key steps of the proof, to substantiate the minimax rate claim.","section":"Theoretical analysis"}],"minor_comments":[{"comment":"Abstract: the distinction between the four scenarios (known/unknown density ratio crossed with bounded/finite-moment) is central but is described only at a high level; a single clarifying sentence on the data requirements for each case would improve readability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and positive assessment of the manuscript's contributions. The request for explicit details on the approximation theorem is reasonable and will be addressed by expanding the theoretical section in the revision.","responses":[{"response":"We agree that the approximation theorem is central to the claims and that its current presentation would benefit from greater explicitness. In the revised manuscript, we will add a dedicated subsection that states the theorem in full, specifying: (i) the Holder smoothness indices α for the regression function and β for the density ratio; (ii) the precise dependence of network depth L and width W on these indices and dimension d (with constants scaling as O(d^C) for some C independent of d); (iii) the function classes (e.g., bounded or finite-moment density ratios, Sobolev-type balls for the regression function). We will also outline the key proof steps: first, a univariate ReLU approximation lemma with polynomial constants; second, a tensor-product construction that preserves the polynomial scaling in d; third, an error decomposition separating approximation, estimation, and density-ratio estimation errors. These additions will directly support the minimax optimality and the mitigation of the curse of dimensionality.","revision_made":"yes","referee_comment":"[Theoretical analysis] Theoretical analysis section: the novel approximation result asserting polynomial (rather than exponential) dependence of network constants on dimension is load-bearing for both the sharper non-asymptotic stochastic error bounds and the claimed mitigation of the curse of dimensionality. The precise statement of the theorem (including the dependence on smoothness indices, network depth/width, and the precise function classes for the regression function and density ratio) must be given explicitly, together with the key steps of the proof, to substantiate the minimax rate claim."}],"tokens_in":1350,"tokens_out":389,"duration_ms":13209,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this work puts repeated measurements regression together with covariate shift correction via density ratios, estimates both the ratio and the target function with ReLU feedforward nets, and supplies a new approximation theory where the constants scale polynomially in dimension instead of exponentially.\n\nThey derive non-asymptotic error bounds for both the known-ratio and unknown-ratio cases, covering bounded ratios and the finite-moment unbounded case. The bounds are claimed to be minimax optimal under the repeated measurements model, and the polynomial scaling is used to sharpen the stochastic error term. The abstract and stress-test note give no sign of internal contradictions or hidden exponential factors in the rate derivations.\n\nWhat they do well is keep the framework explicit: separate handling for known versus estimated ratios, plus the usual simulation and real-data check. The assumption that the target function and density ratio lie in ReLU-approximable classes under stated smoothness is standard for this style of work and is stated up front.\n\nThe soft spot is that everything rests on those approximation classes and on the details of how the polynomial constants are obtained; without the full proofs it is impossible to verify the constants or the exact minimax claim. That is a normal limitation for a theory paper at this stage rather than a fatal gap.\n\nThis is for people who care about nonparametric rates under transfer and repeated observations. The thinking looks clear and the literature engagement appears honest, so the paper deserves a serious referee even if the final rates need tightening.","headline":"The paper's real contribution is a ReLU approximation result with polynomial dimension dependence that lets them get non-asymptotic minimax rates for density-ratio transfer learning in repeated-measures regression.","tokens_in":2350,"tokens_out":383,"would_cite":false,"duration_ms":21449,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ReLU networks estimate target regression under covariate shift by weighting with the density ratio and achieve minimax optimal rates via polynomial-dimension approximation bounds.","keywords":["nonparametric regression","covariate shift","repeated measurements","transfer learning","ReLU neural networks","density ratio estimation","minimax optimal rates","approximation theory"],"falsifier":"Numerical experiments in which the observed convergence rate of the estimator is slower than the stated minimax rate, or in which the network approximation error constants are observed to grow exponentially rather than polynomially with dimension.","tokens_in":2666,"feed_emoji":"📈","tokens_out":720,"duration_ms":19337,"temperature":0.7,"pith_summary":"The paper addresses nonparametric regression when responses cannot be observed in the target domain but are available in a source domain whose covariates follow a different distribution. It transfers information by reweighting source observations with the density ratio between the two covariate distributions, then fits the target regression function with ReLU feedforward neural networks. Both the known-ratio and unknown-ratio cases are treated, along with bounded and finite-moment ratio conditions. The central technical step is a new approximation result showing that the constants governing network size and error grow only polynomially with dimension rather than exponentially, which removes the dominant source of the curse of dimensionality and produces non-asymptotic bounds that match the minimax rate for repeated-measurements designs.","feed_headline":"ReLU nets hit minimax rates for regression under covariate shift","feed_subtitle":"Density-ratio correction plus a new polynomial-dimension approximation bound yields optimal convergence for repeated-measurements data.","key_machinery":"Density-ratio reweighting to correct covariate shift, paired with ReLU feedforward neural network approximation of both the regression function and (when unknown) the density ratio, where the approximation constants depend polynomially on dimension.","core_discovery":"Under the repeated-measurements setting, the proposed ReLU-FNN estimators of the target regression function, after density-ratio correction, attain the minimax optimal convergence rate; the proof relies on a novel approximation theory in which the constants that appear in the network-parameter bounds depend polynomially, rather than exponentially, on the input dimension.","pith_inferences":["The polynomial dependence on dimension may make the method viable in moderately high-dimensional covariate settings where exponential constants would have rendered neural-network approximation impractical.","The same approximation technique could be applied to other transfer-learning or distribution-shift problems whose error analysis is currently limited by exponential dimension dependence.","Testing the method on data sets with known repeated-measurements structure would provide a direct check on whether the claimed rate improvement materializes in finite samples."],"forward_implications":["The same polynomial-dimension approximation theory supplies sharper stochastic-error bounds for both the known-ratio and unknown-ratio estimators.","The estimators remain consistent and rate-optimal when the density ratio is unbounded but satisfies only finite-moment conditions.","Separate procedures are given for the case in which the density ratio is known versus the case in which it must be estimated jointly with the regression function.","The approach directly accommodates the repeated-measurements structure in the source domain."],"fun_headline_variants":["ReLU FNNs attain minimax rates for regression with repeated measurements","Density ratio correction enables minimax optimal convergence in covariate shift","Polynomial dependence on dimension in ReLU network approximation theory","Optimal rates achieved by ReLU nets after density ratio correction"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The target regression function and the density ratio can be approximated to the required accuracy by ReLU neural networks under the given smoothness conditions, and the density ratio is either uniformly bounded or has finite moments.","fun_headline_variants_meta":{"raw":{"variants":["ReLU FNNs attain minimax rates for regression with repeated measurements","Density ratio correction enables minimax optimal convergence in covariate shift","Polynomial dependence on dimension in ReLU network approximation theory","Optimal rates achieved by ReLU nets after density ratio correction"]},"model":"grok-4.3","cost_usd":0.005899,"raw_usage":{"total_tokens":2801,"prompt_tokens":668,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":58987000,"prompt_tokens_details":{"text_tokens":668,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2066,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":668,"tokens_out":67,"duration_ms":18023,"temperature":1.0,"reasoning_tokens":2066,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T23:56:58.399086+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Numerical experiments in which the observed convergence rate of the estimator is slower than the stated minimax rate, or in which the network approximation error constants are observed to grow exponentially rather than polynomially with dimension.","supporting_citations":[],"review_version":1}