Pith. sign in

REVIEW 4 major objections 6 minor 4 references

Using infrared gas sensors in an in-vitro dynamic gut model for detecting short-chain fatty-acids: Technical Report

T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A dual-channel infrared gas sensor and a machine-learning classifier can separate high from low acetate and propionate in an in-vitro gut's headspace; butyrate classification is at chance.

desk verdict Preliminary feasibility report with a genuinely interesting sensor combination, but the headline accuracies are inflated by feature selection before cross-validation and butyrate is at chance; treat as work in progress, not as a measurement claim. read the letter →

arxiv 1909.11177 v1 pith:XCTBVZSD submitted 2019-08-22 physics.ins-det eess.SP

classification physics.ins-deteess.SP
keywords infraredgassensorshort-chainfattyacidsbutyrateelectronicnosemachinelearningclassificationin-vitrogutmodelonlinebioreactormonitoringCO2andhydrocarbonsensing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This technical report asks whether a dual-channel infrared gas sensor placed in the headspace of an in-vitro gut model can track the levels of short-chain fatty acids (SCFAs) produced by fermentation. The authors connected the sensor to two parallel colon vessels fed different diets over four days, and used machine-learning classifiers on features extracted from the CO2 and hydrocarbon signals. They report 17/19 correct high/low classifications for acetate, 14/19 for propionate, and 10/19 for butyrate, meaning only the first two clearly exceed chance. The authors read this as evidence that online, low-cost SCFA monitoring could eventually replace batch sampling and gas chromatography, and that different sensor channels carry complementary information.

What carries the argument

The load-bearing object is a dual-channel infrared gas sensor whose two absorption channels respond to CO2 and to hydrocarbons, mounted in a closed gas loop with the headspace of each gut-model vessel. Because neither channel is chemically specific to acetate, propionate, or butyrate, the measured signals function as indirect proxies for fermentation activity rather than direct SCFA measurements. The signal-processing chain carries the argument: per-reading feature extraction (36 features including standard deviations, amplitude differences, kurtosis, medians, and transform-derived statistics), feature selection for each SCFA, and leave-one-out cross-validation with a classifier that maps features onto high or low concentration relative to the experiment's average. The CO2, hydrocarbon, and reference channels together are what the accuracy figures describe.

What would settle it

Run the same dual-channel sensor against sealed vessels with known SCFA vapor concentrations in nitrogen while varying CO2 and hydrocarbon backgrounds independently; if classification accuracy on the SCFA classes drops to chance once CO2 is held constant or removed, the headspace signal is not directly carrying SCFA information.

Watch

Extended reading notes

Core claim

In a 96-hour experiment with two parallel proximal-colon vessels of an in-vitro gut simulator, one fed a control diet and one a butyrogenic fiber supplement, a dual-channel infrared sensor measuring CO2 and hydrocarbons was connected in a closed loop to the vessel headspaces. Ground-truth acetate, propionate, and butyrate concentrations were obtained by gas chromatography before each of the three daily feedings. From each 55-minute sensor reading the authors extracted 36 features from the reference, CO2, and hydrocarbon channels, selected the best features per substance, and performed leave-one-out cross-validation on 19 labeled readings. The classifier labeled each reading as high or low relative to the four-day average: 17/19 correct for acetate, 14/19 for propionate, and 10/19 for butyrate. The authors conclude that different signals are predictive of different substances—CO2-derived features dominate for acetate, hydrocarbon-derived features for butyrate—and that the approach shows potential for online metabolite detection, while noting that the butyrate result is essentially at chance and that sensor design still needs work.

Load-bearing premise

The load-bearing premise is that the CO2 and hydrocarbon signals measured in the vessel headspace carry information about SCFA levels; the sensor was not calibrated on known SCFA mixtures, so if those signals mostly reflect general fermentation or flushing artifacts, the classification accuracies will not transfer.

Editorial extensions

If this is right

  • If the headspace readings genuinely track SCFA levels, online sensors could reduce the need for manual sampling and gas-chromatography analysis in gut-model experiments, cutting cost and delay.
  • Acetate and propionate can be separated into high and low classes from a single dual-channel IR sensor under the tested conditions, with acetate the easiest (17/19).
  • Butyrate, the compound emphasized for gut health, was not classified beyond chance (10/19), so the current prototype cannot serve as a butyrate monitor.
  • Because different substances rely on different sensor channels, adding more IR channels or sensors targeting other wavelengths could improve per-SCFA discrimination.
  • If implemented online, such monitoring would let experimenters catch failures like the pH-controller error that disrupted one of the two vessels, enabling faster intervention.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the report leaves implicit is that the CO2 and hydrocarbon channels are not shown to respond to SCFA molecules themselves; the accuracies may reflect total fermentation activity rather than individual acid identity, which matters for transfer to real-time monitoring.
  • A direct extension would be to run the same sensor on sealed gas standards with known SCFA concentrations in nitrogen while holding CO2 fixed; if classification collapses when CO2 is controlled, the acetate and propionate signal is likely an indirect CO2 artifact.
  • The chance-level butyrate result suggests that a butyrate monitor will need a sensor with a spectral window where butyrate absorbs, or a dedicated butyrate-selective channel, rather than a broadband hydrocarbon response.
  • The binary high/low split at the experiment average makes the reported accuracies sensitive to the distribution of readings; reanalyzing the same data as regression or with three levels, as the authors mention, would test whether the relationship is graded or threshold-bound.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This technical report describes an exploratory study in which a commercially available dual-channel infrared gas sensor (sensitive to CO2 and hydrocarbons) was connected in a closed loop to the headspace of two SHIME in-vitro gut-model vessels, with the goal of detecting short-chain fatty acids (SCFAs) such as acetate, propionate, and butyrate. Gas readings were collected over 96 hours, yielding 19 measurement windows with GC-MS-based ground-truth SCFA concentrations. From each window, 12 features were extracted from each of three sensor signals (reference, ChA, ChB), giving 36 features per sample. The authors binarized SCFA levels into high and low relative to the mean of the 19 readings, selected features separately for each SCFA on the full data set, and then performed leave-one-out cross-validation. They report accuracies of 17/19 for acetate, 14/19 for propionate, and 10/19 for butyrate. The paper concludes that online detection of SCFAs via IR-sensor signals and machine learning shows potential, while acknowledging a pH-controller malfunction in one vessel that altered fermentation dynamics.

Significance. If the reported accuracies were unbiased, the study would be a useful step toward low-cost online monitoring of SCFAs in bioreactors and gut models, addressing a genuine unmet need in gut-health research. The paper has notable strengths: the ground-truth labels come from independent GC-MS measurements, the experiments were conducted in a realistic SHIME system rather than a synthetic gas rig, and the authors transparently disclose the pH-controller error and the exploratory nature of the analysis. These virtues make the dataset potentially valuable for further study. However, the central empirical claim—that the sensor plus machine learning can classify high versus low SCFA levels—is not yet demonstrated, because the evaluation protocol leaks information from the held-out samples into feature selection, no baseline classifier is reported, and the butyrate result is indistinguishable from chance. The significance of the work therefore rests on whether a re-analysis with a properly nested validation can confirm the accuracy estimates.

major comments (4)
  1. [Section 3, Table 2] The feature-selection procedure is performed on the complete data set before the leave-one-out cross-validation loop. Because the held-out sample's label can influence which features are selected, the reported accuracies (17/19 acetate, 14/19 propionate, 10/19 butyrate) are optimistically biased and do not measure generalization as claimed. With 36 features and only 19 readings, this bias can be substantial. To support the central claim, the authors should use a nested cross-validation scheme in which feature selection is repeated inside each training fold, or use an embedded feature-selection method that does not see held-out labels, and report the selected features' stability across folds.
  2. [Section 3, Table 2] No baseline classifier or null model is reported. With a binary high/low task and 19 samples, a trivial majority-class classifier would already achieve a notable accuracy, and the butyrate result of 10/19 is consistent with random guessing. The authors should report a permutation test, a dummy-classifier comparison, and binomial confidence intervals for the accuracies (or, better, confusion matrices with sensitivity and specificity per class). Additionally, the high/low threshold is defined as the mean of the same 19 readings used for evaluation; the sensitivity of the reported accuracies to this threshold choice should be assessed.
  3. [Section 2.3 and Section 3] The IR sensor measures CO2 and hydrocarbons, not SCFA vapors directly, and no calibration is presented linking the sensor's response to known SCFA concentrations. The statement in Section 2.1 that the sensor 'will only react to the target gas' is not supported by the sensor specifications in Table 1, which list CO2 and hydrocarbons as the responsive gases. Without a calibration experiment using known SCFA mixtures or a gas-chromatographic/mass-balance validation of the headspace signal, it remains possible that the classification is driven by general fermentation CO2 or flushing artifacts rather than by SCFA-specific information. This is load-bearing for the claim that the sensor can detect SCFAs, and it should be addressed with targeted experiments or at least with a clear statement that the reported results concern a proxy signal whose causal link to SCFA concentration is not yet established.
  4. [Section 3, pH-controller error] The disclosed pH-controller error—whereby PC2's pH was regulated using PC1's probe—led to acidification and slowed microbial fermentation in the NAXUS-treated vessel. The paper acknowledges this and notes that a wide range of SCFA concentrations was nonetheless observed, but the analysis treats all 19 readings as if they came from comparable experimental conditions. The authors should report whether the classifier's decisions are confounded by vessel identity (PC1 versus PC2), for example by showing per-vessel accuracy or by including vessel identity as a covariate. If the classifier is simply separating the two vessels, the reported SCFA accuracies would be inflated and the sensor's specificity for individual SCFAs would not be established.
minor comments (6)
  1. [Abstract and Introduction] The abstract and introduction describe 'a novel sensor based on infrared technology,' but the sensor appears to be an off-the-shelf dual-channel IR device; the novelty lies in the application and signal-analysis pipeline, which should be stated more precisely.
  2. [Section 2.1] The text says 'To our knowledge, it is unknown if the target SCFA gases for our application absorb IR light or not,' which is inconsistent with Figure 1 showing NIST infrared spectra for butyrate, acetate, and propionate. Clarify what is unknown: whether the specific LED wavelengths used by the sensor overlap with SCFA absorption bands.
  3. [Section 2.1] Figure references appear to be off by one: Section 2.1 refers to 'Figure 10' and 'Figure 11' for the IR sensor layout and absorption, but later sections use Figures 1–9. Renumber the figures consistently.
  4. [Section 3] There are several typos, including 'CO22 level' (should be 'CO2 level') and 'introduce anti-biotika' (should be 'antibiotics'). These should be corrected.
  5. [Section 3, Table 2] The feature names in Table 2 (refstd, ChA ampdiff, ChA dwtStd, etc.) are not defined in the text. Provide a table or appendix explaining each feature and the signal preprocessing steps.
  6. [Section 3] The description of the machine-learning algorithm is incomplete: the classifier type, hyperparameters, and feature-selection method are not stated, which prevents reproduction. Please specify these details even in a technical-report format.

Circularity Check

2 steps flagged · score 4.0 of 10

Feature selection on all 19 samples before leave-one-out cross-validation, plus high/low labels defined by the full-period mean, make Table 2's accuracies partially circular estimates.

  1. fitted input called prediction [Section 3, Experimental Results, paragraph preceding Table 2]
    "A feature selection algorithm was used to select the best features for each substance. The complete data set contained 19 readings with known ground truth. A full leave-one-out cross-validation was performed and the result is shown in Table 2."

    Feature selection is applied to the complete 19-reading dataset, before the leave-one-out folds are formed. For each held-out reading, its ground-truth label therefore participates in choosing the features on which that same reading is later classified. The Table 2 accuracies (17/19, 14/19, 10/19) are thus not unbiased estimates of generalization: the feature-selection step is fitted on data that includes the held-out sample, so the 'prediction' is partly produced from the target it claims to predict.

  2. self definitional [Section 3, Experimental Results, class definition before Table 2]
    "The level for each substance was categorized into two classes: high level and low level. A high level is defined as above the average concentration over the 4-day period and vice versa."

    The binary target is defined by comparing each reading with the mean of the same 19 readings. The label being predicted is therefore a self-referential function of the full dataset: to know whether a reading is 'high' one must know the average of all readings, including the one being classified. In a real online deployment the 4-day average would not be known in advance, so the reported accuracy describes an in-sample mean-split task rather than prediction against an independently predefined threshold. This does not force the sensor features to be non-predictive, but it makes the stated prediction target partly a construction from the data used for evaluation.

full rationale

The central result is not a pure tautology: the GC-MS labels are independent of the IR sensor signals, so a classifier must still learn a real mapping to exceed chance. However, two evaluation-protocol choices make the reported accuracies partially circular. First, feature selection is performed on the complete dataset before the leave-one-out loop, so the left-out sample's label can influence which features are selected; this is a classic selection-leakage bias and directly affects the numbers in Table 2. Second, the high/low categorization is defined relative to the average concentration computed from the same 19 readings, which makes the target label a function of the full evaluation set rather than an externally fixed threshold. Neither issue alone forces the reported accuracy by construction, and the paper's self-citations (references [3] and [4]) are incidental suggestions for future methods rather than load-bearing evidence. The score of 4 reflects a partial circularity in the evaluation protocol, not a fully self-referential derivation.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central result depends on unvalidated assumptions about sensor-headspace coupling, GC-MS ground truth, and the SHIME model. No new entities are introduced. Two free parameters are dataset-specific choices: the mean-based binary thresholds and the unspecified machine learning feature/architecture choices.

free parameters (2)
  • Binary high/low threshold for each SCFA = Mean acetate, propionate, and butyrate concentration over the 4-day period (numeric values not reported)
    The labels used to train and test the classifier are relative to the same dataset's average, so reported accuracies are specific to this experiment and cannot be compared to fixed physiological cutoffs.
  • Selected feature set (12 of 36 features) and classifier hyperparameters = Not reported
    Feature selection and model tuning are part of the pipeline; without fixed definitions, the reported accuracy is not portable.
assumptions (4)
  • domain assumption Headspace concentrations of CO2 and hydrocarbons measured by the IR sensor are causally or statistically related to SCFA production in the SHIME reactors.
    The sensor is an off-the-shelf CO2/hydrocarbon detector, not a SCFA-specific detector; the paper provides no calibration with known SCFA mixtures (Section 2.3).
  • domain assumption GC-MS measurements of acetate, propionate, and butyrate in the vessels are accurate ground truth.
    The machine learning labels rely entirely on GC-MS; sampling time and representativeness are not validated (Section 2.3).
  • domain assumption The SHIME dynamic gut model reproduces the relevant features of human colonic fermentation.
    The model is cited to Molly et al. [2]; its fidelity to in-vivo SCFA dynamics is taken as given (Section 2.2).
  • ad hoc to paper Binary high/low categorization preserves the information needed for online monitoring.
    The choice of mean-based binary classes discards concentration magnitude and is tailored to this single experiment (Section 3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using infrared gas sensors in an in-vitro dynamic gut model for detecting short-chain fatty-acids: Technical Report." pith.science (2026). https://pith.science/paper/XCTBVZSD

@misc{pith2026190911177,
  author       = {Pith},
  title        = {Pith review of: Using infrared gas sensors in an in-vitro dynamic gut model for detecting short-chain fatty-acids: Technical Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XCTBVZSD}},
  note         = {Machine review of arXiv:1909.11177}
}
read the original abstract

Short-chain fatty acids (SCFAs), including acetate, propionate and butyrate, are organic fatty-acids that are produced when indigested carbohydrates are fermented in the colon by gut-bacteria. Butyrate is especially considered a beneficial compound in relation to gut health and the maintenance of colonic homeostasis. Little is known about butyrate production and the current measurement methods of fecal samples are not representative and there is a strong unmet need to measure SCFAs in-vivo to further understand their role and effect on colonic function in human gut health and disease. The general aim of the project is to develop a novel sensor based on infrared technology that is capable of online detecting short chain fatty acids (SCFA; e.g. acetate, propionate and butyrate) and other small biomolecules (e.g. ammonium) produced in bioreactors. Whereas currently such levels of small biomolecules are usually not detected online due to practical constraints (sampling-analysis-processing of data), reliable online detection could significantly decrease analysis costs. Further, it would allow to perform online quality control of the bioreactors and allow to e.g. detect deviating values at an early stage. Hence, one can react much faster tackling certain problems which will further improve the quality of each experiment. Using machine learning algorithms to process the sensor signals, the results from testing in a so-called artificial gut setup are described in this technical report.

Figures

Figures reproduced from arXiv: 1909.11177 by the authors.

Figure 1
Figure 1. Infrared spectrums for butyrate, acetate, and propionate from the NIST Chemistry WebBook at [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Our sensor setup connected to the SHIME system. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. SHIME gut model at ProDigest, Ghent, Belgium. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Gas response from IR and MOS sensors, with gas flow. During flush action 1, 2, 3 and 4 all SHIME vessels [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Gas response from IR and MOS sensors, with gas flow. [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: A short window of sensor readings of both vessels over a duration of 2 hours. [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Groundtruth of SCFA-levels for a long-time experiment. [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: pH level for a long-time experiment. the best features for each substance. The complete data set contained 19 readings with known ground truth. A full leave-one-out cross-validation was performed and the result is shown in [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 9
Figure 9. Figure 9: Gas response from IR sensors for a long-time experiment. [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages

  1. [1]

    The role of butyrate on colonic function

    Henrike M Hamer, DMAE Jonkers, Koen Venema, SALW Vanhoutvin, FJ Troost, and R-J Brummer. The role of butyrate on colonic function. Alimentary pharmacology & therapeutics, 27(2):104–119, 2008

  2. [2]

    Development of a 5-step multi-chamber reactor as a simulation of the human intestinal microbial ecosystem

    Koen Molly, M Vande Woestyne, and Willy Verstraete. Development of a 5-step multi-chamber reactor as a simulation of the human intestinal microbial ecosystem. Applied microbiology and biotechnology, 39(2):254–258, 1993

  3. [3]

    Fast Classification of Meat Spoilage Markers Using Nanostructured ZnO Thin Films and Unsupervised Feature Learning

    Martin Längkvist, Silvia Coradeschi, Amy Loutfi, and John Bosco Balaguru Rayappan. Fast Classification of Meat Spoilage Markers Using Nanostructured ZnO Thin Films and Unsupervised Feature Learning. Sensors, 13(2):1578–1592, 2013. doi:10.3390/s130201578

  4. [4]

    Unsupervised feature learning for electronic nose data applied to bacteria identification in blood

    Martin Längkvist and Amy Loutfi. Unsupervised feature learning for electronic nose data applied to bacteria identification in blood. In NIPS workshop on Deep Learning and Unsupervised Feature Learning , 2011. 8

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.