REVIEW 3 major objections 2 minor 2 references
Probing the faint end of simulated galaxy counts at z>3
T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read Hydrodynamical simulations underproduce faint compact galaxies at z>3, causing a shortfall in near-infrared counts relative to CANDELS data
desk verdict Multi-field confirmation of the z>3 count tension plus structural evidence for missing compact galaxies, but the FORECAST pipeline's morphology-dependent fidelity remains the untested link. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
FORECAST forward-modeling applied to light-cone catalogs from TNG100 and EAGLE to create mock CANDELS images, followed by direct comparison of detected counts and galaxy structural properties such as compactness and surface brightness.
What would settle it
A new simulation run with updated star-formation or feedback physics that produces matching counts and the observed fraction of compact faint galaxies at z>3 in the same mock images.
Extended reading notes
Core claim
The faint-end discrepancy in near-infrared galaxy counts at z>3 between CANDELS observations and current hydrodynamical simulations arises both from detection losses of diffuse galaxies and from the simulations' inability to produce enough faint compact galaxies with bright central cores.
Load-bearing premise
The forward-modeling pipeline accurately reproduces the observational selection function and completeness limits of the CANDELS data without unaccounted systematics.
Editorial extensions
If this is right
- The deficit is present in all CANDELS fields and emerges at z>3 in both simulations.
- Completeness-corrected observed counts already exceed the intrinsic simulation counts at the 50 percent completeness limit.
- Deeper mock images recover counts near the peak but overpredict the faintest sources.
- Simulations favor diffuse low-surface-brightness systems while underproducing the compact galaxies seen in observations.
Reading between the lines
- The result implies that revisions to how simulations treat early star formation in dense regions or dust attenuation could alter the predicted compact population.
- Deeper or wider surveys at similar redshifts could test whether the observed compact fraction holds beyond current selection effects.
- Alternative simulation suites with different subgrid prescriptions might produce different levels of tension with the same observational dataset.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses the FORECAST forward-modeling code to generate mock CANDELS images from ten light-cone realizations of the TNG100 and EAGLE simulations. It compares both intrinsic catalogs and detected sources in the mocks to observational counts across CANDELS fields, finding a persistent faint-end deficit at z>3 that exceeds completeness-corrected observations already at the 50% limit. Structural analysis indicates simulations underproduce compact galaxies with bright cores while overproducing diffuse systems; the conclusion is that the tension stems from both detection losses and intrinsic limitations in the hydrodynamical simulations' ability to form faint compact galaxies at high redshift.
Significance. If the forward-modeling accurately reproduces CANDELS selection functions without differential systematics between compact and diffuse sources, the result would highlight a concrete shortfall in current simulations for z>3 galaxy populations, pointing to needed improvements in star-formation, feedback, and dust prescriptions. The multi-simulation, multi-field approach and direct intrinsic-vs-mock comparison provide a clear test framework.
major comments (3)
- [Abstract] Abstract and validation description: the pipeline is validated via stellar-mass functions, multi-band photometry, and depth tests, yet no external cross-check (e.g., independent completeness maps or real-source injection recoveries) is reported to confirm that morphology-dependent detection efficiency for compact vs. diffuse z>3 sources matches between mocks and data. This assumption is load-bearing for attributing the tension to simulation physics rather than forward-modeling artifacts.
- [Abstract] Completeness analysis: the claim that GOODS-South completeness-corrected counts already exceed intrinsic simulation counts at the 50% completeness limit lacks reported quantitative error budgets, statistical significance tests, or explicit derivation of the completeness corrections, preventing full evaluation of whether the deficit is robust.
- [Abstract] Structural discrepancy conclusion: the finding that simulations favor diffuse low-surface-brightness systems over compact galaxies with bright central cores depends on the forward-modeling step accurately mapping intrinsic properties to observed detections; without additional validation against differential biases, this risks circularity with the modeling assumptions.
minor comments (2)
- Specify how the ten independent light-cone realizations were generated and how field-to-field variations were quantified.
- Clarify the exact criteria used to classify galaxies as 'compact' versus 'diffuse' in the structural analysis.
Simulated Author's Rebuttal
We thank the referee for the constructive and detailed report. We address each major comment below, agreeing where additional material or clarification is warranted and explaining our position on the analysis. Revisions will be made to improve transparency and robustness.
read point-by-point responses
-
Referee: [Abstract] Abstract and validation description: the pipeline is validated via stellar-mass functions, multi-band photometry, and depth tests, yet no external cross-check (e.g., independent completeness maps or real-source injection recoveries) is reported to confirm that morphology-dependent detection efficiency for compact vs. diffuse z>3 sources matches between mocks and data. This assumption is load-bearing for attributing the tension to simulation physics rather than forward-modeling artifacts.
Authors: Our validations via stellar mass functions, multi-band photometry, and depth tests confirm that the overall selection function and photometric properties are reproduced at the population level. We acknowledge that these do not directly test morphology-dependent completeness for compact versus diffuse sources at z>3. We will add source-injection experiments in the mocks that separately target compact (high central surface brightness) and diffuse morphologies at z>3, reporting recovery fractions as a function of morphology. This will be included in a new subsection of the methods and results. revision: yes
-
Referee: [Abstract] Completeness analysis: the claim that GOODS-South completeness-corrected counts already exceed intrinsic simulation counts at the 50% completeness limit lacks reported quantitative error budgets, statistical significance tests, or explicit derivation of the completeness corrections, preventing full evaluation of whether the deficit is robust.
Authors: The completeness corrections are computed from the recovery rate of simulated galaxies in the mock images as a function of apparent magnitude and redshift, using the ten independent light cones. We will expand the methods section to provide the explicit derivation (including the functional form and binning), report the full error budget (Poisson plus field-to-field variance), and include statistical significance tests (e.g., Kolmogorov-Smirnov or chi-squared) comparing the completeness-corrected observations to the intrinsic simulation counts at the 50% limit. These additions will allow quantitative evaluation of the deficit. revision: yes
-
Referee: [Abstract] Structural discrepancy conclusion: the finding that simulations favor diffuse low-surface-brightness systems over compact galaxies with bright central cores depends on the forward-modeling step accurately mapping intrinsic properties to observed detections; without additional validation against differential biases, this risks circularity with the modeling assumptions.
Authors: The preference for diffuse systems is already visible in the intrinsic light-cone catalogs (prior to image simulation), where the simulations produce fewer objects with small effective radii and high central surface brightness. The forward modeling then applies the same observational effects to both populations. To address potential circularity, we will add a direct comparison of intrinsic structural parameters between simulations and the subset of observed galaxies with reliable size measurements, plus sensitivity tests varying the assumed light-profile parameters in FORECAST. These will be presented in a revised structural analysis section. revision: partial
Circularity Check
No significant circularity; independent comparison of simulations and observations
full rationale
The paper performs direct count and structural comparisons between forward-modeled mocks from public TNG100/EAGLE simulations and public CANDELS catalogs via the FORECAST pipeline. No equations reduce any prediction or result to a fitted parameter or self-defined quantity within the paper. The reference to prior work identifies the initial discrepancy for context but is not load-bearing; the current conclusions rest on new multi-field, depth, and morphology tests against external data. The analysis is self-contained against independent benchmarks with no self-definitional, fitted-input, or uniqueness-imported circular steps.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Probing the faint end of simulated galaxy counts at z>3." pith.science (2026). https://pith.science/paper/2J53KRQC
@misc{pith2026260515893,
author = {Pith},
title = {Pith review of: Probing the faint end of simulated galaxy counts at z>3},
year = {2026},
howpublished = {\url{https://pith.science/paper/2J53KRQC}},
note = {Machine review of arXiv:2605.15893}
}
read the original abstract
Simulations and observations now probe comparable redshift regimes with unprecedented accuracy, enabling direct consistency tests through forward modeling. In a previous work, we identified a faint-end discrepancy between observed and simulated near-infrared galaxy counts in CANDELS GOODS-South. Here we investigate whether this tension originates from the forward-modeling procedure or from limitations of the underlying simulations, and we characterize the galaxy populations responsible for the tension. Using the FORECAST forward-modeling code, we generated ten independent light-cone realizations and mock CANDELS images from the TNG100 and EAGLE simulations. We compared both the intrinsic light-cone catalogs and the mock-image detections with observations, testing dependencies on field and redshift, and validating the pipeline through stellar mass and multi-band analyses. The faint-end deficit is present in all CANDELS fields and appears at z>3 in both simulations. GOODS-South counts corrected for completeness exceed intrinsic simulation counts already at the 50% completeness limit, indicating that the missing population is not simply hidden below the detection threshold. Increasing the depth of mock images recovers the counts near the peak but overpredicts the faintest sources, showing that depth alone cannot resolve the discrepancy. Structural analyses reveal that compact galaxies with bright central cores observed in GOODS-South are underproduced in simulations, which instead favor diffuse low-surface-brightness systems. We conclude that the discrepancy arises both from detection losses of diffuse galaxies and, more fundamentally, from the inability of current hydrodynamical simulations to produce enough faint compact galaxies at z>3. This tension points to the need for improved modeling of early star formation, feedback, and dust treatment.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
— divided by their respective comoving volumes,V TNG = (110.7)3 cMpc3 andV EAGLE =(100) 3 cMpc3. In Fig. A.1 we compare the stellar mass density of the full TNG100 snapshot (golden dots) and the EAGLE snapshot (green dots) with the average mass density of 200 realizations of the corresponding FORECAST partition (purple squares for TNG100 and orange square...
work page 2022
-
[2]
SED-fitting software. We post-processed the noiseless mock images of each band to match the 5σdepth of the corresponding observations (see Table 1 in Sect.2.2.3). Detection was performed independently on the F105W band for both real and mock images, since it is not used as a detection band in the official CANDELS catalogs. For theHand F277W bands, instead...
work page 2021
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.