Pith. sign in

REVIEW 1 cited by

InterFeat: A Pipeline for Finding Interesting Scientific Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13534 v2 pith:R7BTFYCS submitted 2025-05-18 q-bio.QM cs.AIcs.CLcs.IR

classification q-bio.QMcs.AIcs.CLcs.IR
keywords interestingpipelinecandidatesdatadiscoveryfindinginterestingnessinterfeat
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Finding interesting phenomena is the core of scientific discovery, but it is a manual, ill-defined concept. We present an integrative pipeline for automating the discovery of interesting simple hypotheses (feature-target relations with effect direction and a potential underlying mechanism) in structured biomedical data. The pipeline combines machine learning, knowledge graphs, literature search and Large Language Models. We formalize "interestingness" as a combination of novelty, utility and plausibility. On 8 major diseases from the UK Biobank, our pipeline consistently recovers risk factors years before their appearance in the literature. 40--53% of our top candidates were validated as interesting, compared to 0--7% for a SHAP-based baseline. Overall, 28% of 109 candidates were interesting to medical experts. The pipeline addresses the challenge of operationalizing "interestingness" scalably and for any target. We release data and code: https://github.com/LinialLab/InterFeat

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interestingness First Classifiers

    cs.LG 2025-08 conditional novelty 6.0 of 10

    EUREKA uses LLM pairwise comparisons to rank features by interestingness and trains logistic regression on the top-ranked features, producing non-obvious yet above-chance classifiers on six tabular datasets.

Pith tools