Pith. sign in

REVIEW 3 major objections 2 minor 2 references

Probing the faint end of simulated galaxy counts at z>3

T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Hydrodynamical simulations underproduce faint compact galaxies at z>3, causing a shortfall in near-infrared counts relative to CANDELS data

desk verdict Multi-field confirmation of the z>3 count tension plus structural evidence for missing compact galaxies, but the FORECAST pipeline's morphology-dependent fidelity remains the untested link. read the letter →

arxiv 2605.15893 v2 pith:2J53KRQC submitted 2026-05-15 astro-ph.CO astro-ph.GAastro-ph.IM

classification astro-ph.COastro-ph.GAastro-ph.IM
keywords galaxycountshigh-redshiftgalaxieshydrodynamicalsimulationsCANDELSfaint-enddiscrepancyforwardmodelingcompactcompleteness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines a previously noted deficit in faint galaxy counts at high redshift by generating mock CANDELS images from TNG100 and EAGLE simulations using forward modeling. It finds the shortfall appears at z>3 across fields, persists after completeness corrections, and cannot be fixed by simply going deeper. Structural comparisons show simulations produce too many diffuse low-surface-brightness galaxies and too few compact ones with bright cores. The work concludes that both observational selection effects and fundamental limitations in how simulations form early galaxies contribute to the mismatch.

What carries the argument

FORECAST forward-modeling applied to light-cone catalogs from TNG100 and EAGLE to create mock CANDELS images, followed by direct comparison of detected counts and galaxy structural properties such as compactness and surface brightness.

What would settle it

A new simulation run with updated star-formation or feedback physics that produces matching counts and the observed fraction of compact faint galaxies at z>3 in the same mock images.

Watch

Extended reading notes

Core claim

The faint-end discrepancy in near-infrared galaxy counts at z>3 between CANDELS observations and current hydrodynamical simulations arises both from detection losses of diffuse galaxies and from the simulations' inability to produce enough faint compact galaxies with bright central cores.

Load-bearing premise

The forward-modeling pipeline accurately reproduces the observational selection function and completeness limits of the CANDELS data without unaccounted systematics.

Editorial extensions

If this is right

  • The deficit is present in all CANDELS fields and emerges at z>3 in both simulations.
  • Completeness-corrected observed counts already exceed the intrinsic simulation counts at the 50 percent completeness limit.
  • Deeper mock images recover counts near the peak but overpredict the faintest sources.
  • Simulations favor diffuse low-surface-brightness systems while underproducing the compact galaxies seen in observations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The result implies that revisions to how simulations treat early star formation in dense regions or dust attenuation could alter the predicted compact population.
  • Deeper or wider surveys at similar redshifts could test whether the observed compact fraction holds beyond current selection effects.
  • Alternative simulation suites with different subgrid prescriptions might produce different levels of tension with the same observational dataset.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper uses the FORECAST forward-modeling code to generate mock CANDELS images from ten light-cone realizations of the TNG100 and EAGLE simulations. It compares both intrinsic catalogs and detected sources in the mocks to observational counts across CANDELS fields, finding a persistent faint-end deficit at z>3 that exceeds completeness-corrected observations already at the 50% limit. Structural analysis indicates simulations underproduce compact galaxies with bright cores while overproducing diffuse systems; the conclusion is that the tension stems from both detection losses and intrinsic limitations in the hydrodynamical simulations' ability to form faint compact galaxies at high redshift.

Significance. If the forward-modeling accurately reproduces CANDELS selection functions without differential systematics between compact and diffuse sources, the result would highlight a concrete shortfall in current simulations for z>3 galaxy populations, pointing to needed improvements in star-formation, feedback, and dust prescriptions. The multi-simulation, multi-field approach and direct intrinsic-vs-mock comparison provide a clear test framework.

major comments (3)
  1. [Abstract] Abstract and validation description: the pipeline is validated via stellar-mass functions, multi-band photometry, and depth tests, yet no external cross-check (e.g., independent completeness maps or real-source injection recoveries) is reported to confirm that morphology-dependent detection efficiency for compact vs. diffuse z>3 sources matches between mocks and data. This assumption is load-bearing for attributing the tension to simulation physics rather than forward-modeling artifacts.
  2. [Abstract] Completeness analysis: the claim that GOODS-South completeness-corrected counts already exceed intrinsic simulation counts at the 50% completeness limit lacks reported quantitative error budgets, statistical significance tests, or explicit derivation of the completeness corrections, preventing full evaluation of whether the deficit is robust.
  3. [Abstract] Structural discrepancy conclusion: the finding that simulations favor diffuse low-surface-brightness systems over compact galaxies with bright central cores depends on the forward-modeling step accurately mapping intrinsic properties to observed detections; without additional validation against differential biases, this risks circularity with the modeling assumptions.
minor comments (2)
  1. Specify how the ten independent light-cone realizations were generated and how field-to-field variations were quantified.
  2. Clarify the exact criteria used to classify galaxies as 'compact' versus 'diffuse' in the structural analysis.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive and detailed report. We address each major comment below, agreeing where additional material or clarification is warranted and explaining our position on the analysis. Revisions will be made to improve transparency and robustness.

read point-by-point responses
  1. Referee: [Abstract] Abstract and validation description: the pipeline is validated via stellar-mass functions, multi-band photometry, and depth tests, yet no external cross-check (e.g., independent completeness maps or real-source injection recoveries) is reported to confirm that morphology-dependent detection efficiency for compact vs. diffuse z>3 sources matches between mocks and data. This assumption is load-bearing for attributing the tension to simulation physics rather than forward-modeling artifacts.

    Authors: Our validations via stellar mass functions, multi-band photometry, and depth tests confirm that the overall selection function and photometric properties are reproduced at the population level. We acknowledge that these do not directly test morphology-dependent completeness for compact versus diffuse sources at z>3. We will add source-injection experiments in the mocks that separately target compact (high central surface brightness) and diffuse morphologies at z>3, reporting recovery fractions as a function of morphology. This will be included in a new subsection of the methods and results. revision: yes

  2. Referee: [Abstract] Completeness analysis: the claim that GOODS-South completeness-corrected counts already exceed intrinsic simulation counts at the 50% completeness limit lacks reported quantitative error budgets, statistical significance tests, or explicit derivation of the completeness corrections, preventing full evaluation of whether the deficit is robust.

    Authors: The completeness corrections are computed from the recovery rate of simulated galaxies in the mock images as a function of apparent magnitude and redshift, using the ten independent light cones. We will expand the methods section to provide the explicit derivation (including the functional form and binning), report the full error budget (Poisson plus field-to-field variance), and include statistical significance tests (e.g., Kolmogorov-Smirnov or chi-squared) comparing the completeness-corrected observations to the intrinsic simulation counts at the 50% limit. These additions will allow quantitative evaluation of the deficit. revision: yes

  3. Referee: [Abstract] Structural discrepancy conclusion: the finding that simulations favor diffuse low-surface-brightness systems over compact galaxies with bright central cores depends on the forward-modeling step accurately mapping intrinsic properties to observed detections; without additional validation against differential biases, this risks circularity with the modeling assumptions.

    Authors: The preference for diffuse systems is already visible in the intrinsic light-cone catalogs (prior to image simulation), where the simulations produce fewer objects with small effective radii and high central surface brightness. The forward modeling then applies the same observational effects to both populations. To address potential circularity, we will add a direct comparison of intrinsic structural parameters between simulations and the subset of observed galaxies with reliable size measurements, plus sensitivity tests varying the assumed light-profile parameters in FORECAST. These will be presented in a revised structural analysis section. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; independent comparison of simulations and observations

full rationale

The paper performs direct count and structural comparisons between forward-modeled mocks from public TNG100/EAGLE simulations and public CANDELS catalogs via the FORECAST pipeline. No equations reduce any prediction or result to a fitted parameter or self-defined quantity within the paper. The reference to prior work identifies the initial discrepancy for context but is not load-bearing; the current conclusions rest on new multi-field, depth, and morphology tests against external data. The analysis is self-contained against independent benchmarks with no self-definitional, fitted-input, or uniqueness-imported circular steps.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No new free parameters, axioms, or invented entities are introduced; the work relies on existing public simulations and observational catalogs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Probing the faint end of simulated galaxy counts at z>3." pith.science (2026). https://pith.science/paper/2J53KRQC

@misc{pith2026260515893,
  author       = {Pith},
  title        = {Pith review of: Probing the faint end of simulated galaxy counts at z>3},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2J53KRQC}},
  note         = {Machine review of arXiv:2605.15893}
}
read the original abstract

Simulations and observations now probe comparable redshift regimes with unprecedented accuracy, enabling direct consistency tests through forward modeling. In a previous work, we identified a faint-end discrepancy between observed and simulated near-infrared galaxy counts in CANDELS GOODS-South. Here we investigate whether this tension originates from the forward-modeling procedure or from limitations of the underlying simulations, and we characterize the galaxy populations responsible for the tension. Using the FORECAST forward-modeling code, we generated ten independent light-cone realizations and mock CANDELS images from the TNG100 and EAGLE simulations. We compared both the intrinsic light-cone catalogs and the mock-image detections with observations, testing dependencies on field and redshift, and validating the pipeline through stellar mass and multi-band analyses. The faint-end deficit is present in all CANDELS fields and appears at z>3 in both simulations. GOODS-South counts corrected for completeness exceed intrinsic simulation counts already at the 50% completeness limit, indicating that the missing population is not simply hidden below the detection threshold. Increasing the depth of mock images recovers the counts near the peak but overpredicts the faintest sources, showing that depth alone cannot resolve the discrepancy. Structural analyses reveal that compact galaxies with bright central cores observed in GOODS-South are underproduced in simulations, which instead favor diffuse low-surface-brightness systems. We conclude that the discrepancy arises both from detection losses of diffuse galaxies and, more fundamentally, from the inability of current hydrodynamical simulations to produce enough faint compact galaxies at z>3. This tension points to the need for improved modeling of early star formation, feedback, and dust treatment.

Figures

Figures reproduced from arXiv: 2605.15893 by the authors.

Figure 1
Figure 1. Number counts in the H band in the CANDELS fields (from left to right: COSMOS, EGS, UDS, GS, GSN). Both the mock detections (purple) and the IU counts (blue) are derived from the five cmd-TNG100 realizations; shaded areas indicate the 1σ scatter across the realizations. The mock images are simulated at the corresponding survey depth (i.e. at the 5σ limiting magnitude of the real datasets). Dashed colored lines indic… view at source ↗
Figure 2
Figure 2. H band number counts divided into redshift bins (from top to bottom: z=0.0–2.9, z=3.0–3.6, and z=3.7–5.0). Left pan￾els show results from TNG100, right panels to EAGLE. Mock datasets, obtained from the five cmd- realizations, are repre￾sented by solid line with 1σ shaded area. The IU (blue for TNG100 and magenta for EAGLE) and the sources detected on the mock image (black for TNG100 and brown for EAGLE) are compared… view at source ↗
Figure 4
Figure 4. Comparison between observed H-band number counts in the CANDELS GS field (dashed bright red) and IU counts from five cmd-TNG100 realizations (blue line with 1σ shaded area). The dashed dark red line shows the completeness￾corrected GS counts, obtained using the completeness curve at log10(Rhl/pix) = 0.6. Vertical lines indicate the 50% completeness magnitudes derived for different source sizes (log10(Rhl/pix) = 0.3,… view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: Completeness curves for fake sources injected into the [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 5
Figure 5. Figure 5: Comparison of H-band number counts between the CAN￾DELS GS field (dashed red) and the FORECAST mock dataset based on five cmd-TNG100 realizations. Mock detections at GS depth are shown in black with 1σ shaded area, while the ‘deeper’ mock (cyan with shaded area) is obt…
Figure 6
Figure 6. Figure 6: Characterization of the sources in GOODS-South (red) and in a mock image created from one of the [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Properties of galaxies at z > 3 detected in mock images with noise reduced by a factor of 10 relative to the nominal depth (‘deeper’ image; kept sources, violet/gray) and missed at nominal depth (lost, orange/pink). The upper panels show the intrinsic Input Universe pr…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages

  1. [1]

    — divided by their respective comoving volumes,V TNG = (110.7)3 cMpc3 andV EAGLE =(100) 3 cMpc3. In Fig. A.1 we compare the stellar mass density of the full TNG100 snapshot (golden dots) and the EAGLE snapshot (green dots) with the average mass density of 200 realizations of the corresponding FORECAST partition (purple squares for TNG100 and orange square...

  2. [2]

    We post-processed the noiseless mock images of each band to match the 5σdepth of the corresponding observations (see Table 1 in Sect.2.2.3)

    SED-fitting software. We post-processed the noiseless mock images of each band to match the 5σdepth of the corresponding observations (see Table 1 in Sect.2.2.3). Detection was performed independently on the F105W band for both real and mock images, since it is not used as a detection band in the official CANDELS catalogs. For theHand F277W bands, instead...

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.