Pith. sign in

REVIEW 4 major objections 5 minor 31 references

Learning Real Estate Automated Valuation Models from Heterogeneous Data Sources

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that a three-feature automated valuation model can reproduce professional real-estate appraisals to within about €21,000 using OMI price bands, surface, and comparable-property prices.

desk verdict Useful applied finding: OMI price bands plus comparables predict Turin appraisals well, but the out-of-sample generalization claim is not yet supported because the paper doesn't document feature timing. read the letter →

arxiv 1909.00704 v1 pith:OYZTDXHB submitted 2019-09-02 cs.CY cs.LGstat.ML

classification cs.CYcs.LGstat.ML
keywords automatedvaluationmodelsrealestateappraisalmachinelearningOMIpricebandscomparablepropertieswebcrawlingopendatafeatureselection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a practical automated valuation model for Italian residential property can be built from heterogeneous web and open data rather than from detailed on-site inspection. Its central claim is that once three numbers are known—the official OMI price band for the property's zone, the surface in square meters, and the average advertised or appraised price per square meter of nearby comparable properties—the remaining structural and neighborhood features add almost nothing. Trained on professional appraisals in Turin and evaluated on a held-out test set, the best regressors reach a root mean square error of about €21,000 with $R^2\approx 0.96$, and about €17,000 on 58 additional properties elsewhere in Italy. If the claim holds, banks and appraisal companies could obtain near-expert valuations in real time over the whole country, with the web crawler supplying the local market signal.

What carries the argument

The load-bearing mechanism is the constructed feature "average price per square meter of comparable properties," combined with the official OMI price band and the property's surface. Comparables are obtained by simulating appraiser practice: a web crawler searches nearby advertised properties and previous appraisals near the target, filters by distance and similarity, and averages their unit prices into a single number. This feature encodes the local market directly, which is why adding it drops test error from about €35,000 to about €21,000. Tree-ensemble regressors—extra trees, random forests, and bagging—are the learning machinery, and their feature-importance rankings were the tool that revealed the OMI-centered feature set.

What would settle it

Re-run the out-of-sample test with date-restricted features: for each of the 58 properties, crawl comparables and fetch OMI bands only from the period on or before the appraisal date, and check whether the error stays near €17,000. The paper's Section 5.3 reports the result but does not describe the comparables pipeline for these properties, so this is where the generalization claim can be tested.

Watch

Extended reading notes

Core claim

The key discovery is that a reduced feature set—OMI minimum and maximum price per square meter for the zone, property surface, and the average unit price of comparable properties—reproduces professional appraisals more accurately than a rich hedonic model with dozens of structural, point-of-interest, and geographic features. Feature-importance analysis using tree ensembles showed the OMI values dominating the ranking, which led the authors to drop the OMI area name, distance from center, and all PoI counts. With the comparables feature added, bagging, extra trees, and random forests all give a mean error around €21,000 and $R^2$ around 0.96 on the Turin test set. The same trained model applied to 58 out-of-sample appraisals in different Italian cities and different years gives a mean error of about €17,000, which the paper presents as evidence that the model is predictive beyond the training city.

Load-bearing premise

The nationwide claim rests on the 58 out-of-sample appraisals being measured with the same price-band and comparable-property data that were available on or before each appraisal's date; if later information leaked into those features, the €17,000 error is too optimistic.

Editorial extensions

If this is right

  • A national Italian automated valuation service can be trained on one city and reused elsewhere, because the OMI band and comparables carry the local market signal.
  • Appraisal companies can use the model as an automated validation tool: the paper reports that the error level of about €21,000 on a mean value of about €189,000 was considered acceptable by the domain experts involved.
  • Lenders could issue real-time draft mortgage valuations from a web crawler plus three features, deferring the full expert visit to later in the process.
  • Collecting hedonic details such as elevator, floor, and maintenance status may no longer be necessary for a first-cut valuation when OMI values are available, reducing the cost of appraisal data acquisition.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the comparables feature is computed from advertised asking prices rather than transaction prices, it likely captures list-price levels; a natural extension is to test whether replacing it with transaction prices from deeds, where publicly available, changes accuracy.
  • The transferability of the approach to other countries depends on finding an administrative analog of OMI price bands; in markets without such public zoning valuations, the model would need a different spatial prior to achieve similar accuracy.
  • The 58-property out-of-sample set is too small to certify a national claim; a stronger test would hold out entire cities or time periods from training.
  • The paper's target is expert appraisal rather than market price, so the model may reproduce systematic deviations in expert judgment; comparing predictions to actual transaction prices would separate these two effects.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes an automated valuation model (AVM) for Italian residential real estate, integrating three data sources: professional appraisal records, open geographic data from the Italian Revenue Agency's OMI observatory and points-of-interest databases, and advertised prices of comparable properties obtained by web crawling. After cleaning, 3983 Turin appraisals are used to train several regressors, with three feature sets: a full hedonic set, a reduced OMI-centered set (OMI min/max price bands plus surface), and the same set augmented with the average price per square meter of comparable properties. The best models achieve an RMSE of roughly 21,000 euros (R2 about 0.96) on a held-out Turin test set, and the paper reports a mean error of about 17,000 euros on 58 out-of-sample properties located in different Italian cities and appraised in 2012-2016. The central claim is that a model trained only on Turin data can be used for real-time valuations across the whole Italian territory.

Significance. If the out-of-sample result is valid, the paper offers a practically relevant result: a small set of features, combining official OMI price bands with web-crawled comparable prices, predicts professional appraisals at an accuracy that domain experts reportedly find acceptable. The use of external benchmarks (OMI bands and advertised comparables) rather than a refit of the target variable reduces, though does not eliminate, circularity concerns, and the use of a separate test set is good practice. The main strengths are the realistic appraisal dataset, the explicit comparison of feature sets, and the attempt to support spatial and temporal generalization with an out-of-sample test. However, the generalization claim currently rests on an underdocumented 58-property experiment, so the significance is conditional on closing that gap.

major comments (4)
  1. [Section 5.3 and Sections 3.2.1/3.3] The nationwide generalization claim rests on the 58-property out-of-sample result, but the paper does not document that the OMI and comparable-property features were computed as of each property's appraisal date. Section 3.2.1 states that OMI price ranges are updated every six months, and Section 5.3 covers appraisals from 2012 to 2016, yet the paper never states which OMI release was used for the 58 properties. Similarly, Section 3.3 describes the comparable crawler for the Turin data but gives no crawl date, query radius, or rule for excluding listings posted after the appraisal date for the out-of-sample set. If these features were obtained from later data, the 17,000-euro error is optimistic and not a prospective measure. This is the most load-bearing gap in the paper.
  2. [Section 4.1 and Sections 5.1-5.2] The feature selection process is not described as a fully independent procedure. Section 4.1 explains that the reduced OMI-centered feature set was adopted after examining feature importances from models trained on the data, and Section 5.2 then reports test-set results for that selected feature set. Because the feature set was chosen using knowledge of the same data (and potentially the same test set), the reported errors may be optimistic. The paper should state explicitly that feature selection was performed only on the training partition, or use a nested/outer holdout procedure, and should report results over multiple random splits to quantify variance.
  3. [Section 5.1, Eq. (3), and Tables 2-4] What the paper calls 'Mean Error' (ME) is in fact the root mean square error: Eq. (3) defines ME as sqrt((1/n) sum (y_i - yhat_i)^2). This is not a mean error in the standard sense (which would be the mean absolute or mean signed error). The terminology propagates to the abstract, Section 5.2, and Section 5.3, where the out-of-sample 'Mean Error of 17,000 Euros' should be called 'RMSE of 17,000 Euros'. The mislabeling makes the reported accuracy appear to be a bias measure rather than an error magnitude and should be corrected throughout.
  4. [Section 5.3 and Conclusions] The claim that the model is 'predictive and practically effective for the whole Italian territory' is stronger than the evidence. The out-of-sample set contains only 58 properties from unspecified Italian cities, and no information is given about their distribution across cities, market segments, or time periods. Even with perfect temporal alignment, 58 properties are a thin basis for a nationwide applicability claim, and the paper should either add details about the composition of this set or temper the conclusion to a preliminary indication.
minor comments (5)
  1. [Section 3.1 vs Section 5.3] The appraisal years are stated inconsistently: Section 3.1 says the original data set contains valuations performed between 2011 and 2016, while Section 5.3 says the Turin valuations used to train the model were performed between 2015 and 2016. Please clarify which years were used after cleaning and how the discrepancy arises.
  2. [Section 4.2] The reported partition sizes (2789 + 596 + 596 = 3981) do not sum to the stated 3983 appraisals. Please check the counts and explain how the 70/15/15 split was applied.
  3. [Tables 2-4] The units of MSE should be stated explicitly. Since ME is expressed in thousands of euros, MSE is in (thousands of euros)^2, which is an unusual unit and may confuse readers.
  4. [Section 5.2] The statement that an error around 21,000 euros has been 'considered acceptable by the domain experts' would be more convincing if the acceptability criterion were defined, for example as a percentage of mean property value or as a comparison with expert disagreement between appraisers.
  5. [Section 2 and Section 5] The claim of outperforming the state of the art in [16] is not supported by a direct comparison table or by re-evaluation of the same benchmark data. Please either include such a comparison or rephrase the claim as an indicative comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the AVM's inputs (OMI price bands, surface, comparable average prices) are externally sourced and are not defined in terms of the target appraisals.

full rationale

The derivation chain is self-contained and non-circular. The target variable is the professional appraisal value (Section 3.1, Table 1), while the winning feature set consists of OMI min/max price bands from the Italian Revenue Agency (Section 3.2.1), surface, and the average price per square meter of comparable properties obtained from previous appraisals or web advertisements (Section 3.3). None of these features is constructed from the target property's own valuation, and no fitted parameter is renamed as a prediction: the regressors are trained on 70% of the Turin appraisals and evaluated on a separate test set plus 58 out-of-sample properties (Sections 4.2, 5.2, 5.3). The paper does cite [16] as a baseline it outperforms, but that citation is external and not load-bearing for any derivation. The main caveats are correctness risks, not circularity: feature selection and hyperparameters were chosen on the same corpus, and Section 5.3 does not document that OMI bands and comparable listings were temporally aligned with each out-of-sample appraisal date. Those issues concern external validity and possible look-ahead leakage, but they do not make the claimed prediction equivalent to its inputs by construction.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central result depends on several hand-chosen preprocessing thresholds, distance weighting choices, and the decision to let OMI price bands dominate prediction. These are not derived from theory, and the most important one, comparable retrieval, is not specified precisely enough to audit.

free parameters (5)
  • PoI threshold distance r = 1 km (initial)
    Chosen, not learned. Controls which nearby points count for each PoI feature (Equation 2, Section 3.2.2).
  • PoI ring weights = 1, 0.5, 0.25, 0.125
    Ad hoc descending weights for four distance rings in Equation 2, with no sensitivity analysis.
  • Data cleaning thresholds = residential only, surface <= 250 m2, valuation 20,000-700,000 euros
    Post hoc exclusions in Section 3.1 that reduce 7,988 to 3,983 appraisals and shape the reported accuracy.
  • Comparable retrieval parameters = unspecified
    The crawling procedure is described only in prose (Section 3.3); the number of comparables, distance threshold, and matching filters are not reported.
  • Algorithm hyperparameters = not enumerated
    K, tree counts, and depths were tuned on the validation set (Section 4.2) but not listed, so the reported results cannot be recreated.
assumptions (5)
  • domain assumption Expert appraisals in the dataset are a valid ground truth for property value.
    The target variable is the appraiser's valuation, not a transaction price, and the model is judged against it (Section 3.1).
  • domain assumption OMI price bands are reliable, homogeneous indicators of market value within each area.
    OMI min and max are used as the dominant features (Section 3.2.1), which assumes the official observatory's ranges are accurate across Italian municipalities.
  • domain assumption Advertised prices of comparable properties are a useful upper-bound proxy for market value.
    The comparables feature averages advertised or previously appraised prices per square meter (Section 3.3).
  • ad hoc to paper A random shuffle split is sufficient to avoid temporal and spatial leakage.
    The data is split 70/15/15 after shuffling (Section 4.2), which ignores the 2011-2016 time span and the spatial clustering of Turin properties; comparables might come from the training period.
  • ad hoc to paper The 58 out-of-sample properties are representative of the whole Italian territory.
    The claim of whole-Italy effectiveness rests on this small convenience sample (Section 5.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning Real Estate Automated Valuation Models from Heterogeneous Data Sources." pith.science (2026). https://pith.science/paper/OYZTDXHB

@misc{pith2026190900704,
  author       = {Pith},
  title        = {Pith review of: Learning Real Estate Automated Valuation Models from Heterogeneous Data Sources},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OYZTDXHB}},
  note         = {Machine review of arXiv:1909.00704}
}
read the original abstract

Real estate appraisal is a complex and important task, that can be made more precise and faster with the help of automated valuation tools. Usually the value of some property is determined by taking into account both structural and geographical characteristics. However, while geographical information is easily found, obtaining significant structural information requires the intervention of a real estate expert, a professional appraiser. In this paper we propose a Web data acquisition methodology, and a Machine Learning model, that can be used to automatically evaluate real estate properties. This method uses data from previous appraisal documents, from the advertised prices of similar properties found via Web crawling, and from open data describing the characteristics of a corresponding geographical area. We describe a case study, applicable to the whole Italian territory, and initially trained on a data set of individual homes located in the city of Turin, and analyze prediction and practical applicability.

Figures

Figures reproduced from arXiv: 1909.00704 by the authors.

Figure 1
Figure 1. Valuation distribution in the Appraisal Data Set [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Graphical representation of the Turin real estate data set divided by OMI area - average prices [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Graphical representation of the Turin real estate data set divided by OMI area - predicted prices [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Relative importance of the Hedonic feature set, using the Random Forest regressor (best 9) [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Relative importance of the Hedonic feature set, using the Gradient Boosting regressor (best 9) [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Results for OMI-centered features (random forest) [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Results for OMI-centered features plus comparables (random forest) [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Out of sample test (prices in multiples of 1,000 [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

31 extracted references · 30 canonical work pages

  1. [1]

    Big data in real estate? from manual appraisal to automated valuation

    Nils Kok, Eija-Leena Koponen, and Carmen Adriana Martínez-Barbosa. Big data in real estate? from manual appraisal to automated valuation. The Journal of Portfolio Management, 43(6):202–211, 2017

  2. [2]

    Data mining: An empirical application in real estate valuation

    Ruben D Jaen. Data mining: An empirical application in real estate valuation. In FLAIRS Conference, pages 314–317, 2002

  3. [3]

    Downie, Gill

    Mary Lou. Downie, Gill. Robson, and Council of Mortgage Lenders. Automated valuation models : an interna- tional perspective. CML Research, London, 2007

  4. [4]

    Rey Carmona

    Julia Núñez Tabales, José Caridad, and F.J. Rey Carmona. Artificial neural networks for predicting real estate price. Revista de Metodos Cuantitativos para la Economia y la Empresa, 15:29–44, 2013

  5. [5]

    A data mining model by using ann for predicting real estate market: Comparative study

    Itedal Sabri Hashim Bahia. A data mining model by using ann for predicting real estate market: Comparative study. International Journal of Intelligence Science, 3(04):162, 2013

  6. [6]

    Applied machine learning project 4 prediction of real estate property prices in montreal, 2014

    Nissan Pow, Emil Janulewicz, and L Liu. Applied machine learning project 4 prediction of real estate property prices in montreal, 2014

  7. [7]

    Machine learning for a london housing price prediction mobile application

    Aaron Ng and Marc Deisenroth. Machine learning for a london housing price prediction mobile application. Technical report, Technical Report, June 2015, Imperial College, London, UK, 2015

  8. [8]

    An extended semi-supervised regression approach with co-training and geographical weighted regression: A case study of housing prices in beijing

    Yi Yang, Jiping Liu, Shenghua Xu, and Yangyang Zhao. An extended semi-supervised regression approach with co-training and geographical weighted regression: A case study of housing prices in beijing. ISPRS International Journal of Geo-Information, 5(1):4, 2016. 12 A PREPRINT - SEPTEMBER 4, 2019

Show all 31 references
  1. [9]

    Urban data streams and machine learning: A case of swiss real estate market

    Vahid Moosavi. Urban data streams and machine learning: A case of swiss real estate market. arXiv preprint arXiv:1704.04979, 2017

  2. [10]

    Goodman and Thomas Thibodeau

    Allen C. Goodman and Thomas Thibodeau. Housing market segmentation and hedonic prediction accuracy. Journal of Housing Economics, 12(3):181–201, 2003

  3. [11]

    Multivariate regression modeling for home value estimates with evaluation using maximum information coefficient

    Gongzhu Hu, Jinping Wang, and Wenying Feng. Multivariate regression modeling for home value estimates with evaluation using maximum information coefficient. In Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing 2012, pages 69–81. Springer, 2013

  4. [12]

    Estimation of a hedonic house price model with bargaining: Evidence from the italian housing market

    Gaetano Lisi and Mauro Iacobini. Estimation of a hedonic house price model with bargaining: Evidence from the italian housing market. 2012

  5. [13]

    The potential of big housing data: An application to the italian real-estate market

    Michele Loberto, Andrea Luciani, and Marco Pangallo. The potential of big housing data: An application to the italian real-estate market. Bank of Italy, Temi di Discussione (Working Paper) No. 1171, 2018

  6. [14]

    Predicting housing value: A comparison of multiple regression analysis and artificial neural networks

    Nguyen Nghiep and Cripps Al. Predicting housing value: A comparison of multiple regression analysis and artificial neural networks. Journal of real estate research, 22(3):313–336, 2001

  7. [15]

    Discovering the hidden structure of house prices with a non-parametric latent manifold model

    Sumit Chopra, Trivikraman Thampy, John Leahy, Andrew Caplin, and Yann LeCun. Discovering the hidden structure of house prices with a non-parametric latent manifold model. In Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining, pag...

  8. [16]

    The economic value of neighborhoods: Predicting real estate prices from the urban environment

    Marco De Nadai and Bruno Lepri. The economic value of neighborhoods: Predicting real estate prices from the urban environment. In Proceedings of 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA), Turin, Italy. IEEE, 2018

  9. [17]

    Machine learning and the spatial structure of house prices and housing returns

    Andrew Caplin, Sumit Chopra, John Leahy, Yann LeCun, and Trivikraman Thampy. Machine learning and the spatial structure of house prices and housing returns. 2008

  10. [18]

    Identifying real estate opportunities using machine learning

    Alejandro Baldominos, Iván Blanco, Antonio José Moreno, Rubén Iturrarte, Oscar Bernárdez, and Carlos Afonso. Identifying real estate opportunities using machine learning. MDPI Applied Sciences, 8(11), 2018

  11. [19]

    Guangliang Gao, Zhifeng Bao, Jie Cao, A. K. Qin, Timos Sellis, and Zhiang Wu. Location-centered house price prediction: A multi-task learning approach. CoRR, abs/1901.01774, 2019

  12. [20]

    House price prediction: hedonic price model vs

    Visit Limsombunchai. House price prediction: hedonic price model vs. artificial neural network. In New Zealand Agricultural and Resource Economics Society Conference, pages 25–26, 2004

  13. [21]

    Training a multilayer perceptron to predict the final selling price of an apartment in co-operative housing society sold in stockholm city with features stemming from open data

    Rasmus Tibell. Training a multilayer perceptron to predict the final selling price of an apartment in co-operative housing society sold in stockholm city with features stemming from open data. In Second Level Master´s Thesis, Stockholm, Sweden. 2015

  14. [22]

    Estimating the performance of random forest versus multiple regression for predicting prices of the apartments

    Marjan ˇCeh, Milan Kilibarda, Anka Lisec, and Branislav Bajat. Estimating the performance of random forest versus multiple regression for predicting prices of the apartments. ISPRS Int. J. Geo-Inf (MDPI), 7(5), 2018

  15. [23]

    House price estimation in hanoi using artificial neural network and support vector machine: Considering effects of status and house quality

    Quang-Thanh Bui, Nhu-Hiep Do, and Hoang Phe. House price estimation in hanoi using artificial neural network and support vector machine: Considering effects of status and house quality. In Proceedings of FIG Working Week 2017, Surveying the world of tomorrow - From digitalizati...

  16. [24]

    Extremely randomized trees

    Pierre Geurts, Damien Ernst, and Louis Wehenkel. Extremely randomized trees. Machine learning, 63(1):3–42, 2006

  17. [25]

    Days on market: Measuring liquidity in real estate markets

    Hengshu Zhu, Hui Xiong, Fangshuang Tang, Qi Liu, Yong Ge, Enhong Chen, and Yanjie Fu. Days on market: Measuring liquidity in real estate markets. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 393–402. ACM, 2016

  18. [26]

    House price estimation from visual and textual features

    Eman Ahmed and Mohamed Moustafa. House price estimation from visual and textual features. In Proc. 8th Int. Joint Conf. on Computational Intelligence (IJCCI), pages 62–68, 2016

  19. [27]

    Open data from the italian public administrations, 2019

    Agid (Agenzia per l’Italia Digitale). Open data from the italian public administrations, 2019

  20. [28]

    Manuale della banca dati dell’osservatorio del mercato immobiliare, 2017

    Osservatorio Mercato Immobiliare. Manuale della banca dati dell’osservatorio del mercato immobiliare, 2017

  21. [29]

    Automated valuation model: An application to the public housing resale market in singapore

    Muhammad Faishal Ibrahim, Fook Jam Cheng, and Kheng How Eng. Automated valuation model: An application to the public housing resale market in singapore. Property Management, 23:357–373, 12 2005

  22. [30]

    The value spatial component in the real estate market: the turin case study

    Elena Fregonara, Diana Rolando, and Patrizia Semeraro. The value spatial component in the real estate market: the turin case study. Aestimum, (60):85, 2012

  23. [31]

    Learning to rank social bots

    Diego Perna and Andrea Tagarelli. Learning to rank social bots. In Proceedings of the 29th on Hypertext and Social Media, HT ’18, pages 183–191, New York, NY , USA, 2018. ACM. 13 A PREPRINT - SEPTEMBER 4, 2019 List of Figures 1 Valuation distribution in the Appraisal Data Set ...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.