{"id":"d0056a0c-12fb-48d1-8246-145360bb476a","arxiv_id":"2608.01544","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A CLIP-based image polarity score predicts Moscow apartment listing prices and selling speed beyond standard property characteristics.","lead":"This paper introduces a simple image-scoring method that uses a pre-trained CLIP model to rate how luxurious or rundown a property looks from photos, then shows the score predicts listing prices and time on market for Moscow apartments. A generalist should care because it offers a cheap, reproducible way to turn visual quality into a numeric variable for economics and market analysis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'objective quality' reading is undermined by an uncontrolled marketing channel: 83% of ads come from professional agents/agencies, yet Table 5 never controls for author type, so the CLIP Q-score may proxy for staging/listing effort rather than physical condition.","rationale":"In good faith, the paper's predictive evidence is genuinely supportive: the CLIP Q-score is computed from a frozen model, is not fitted to prices, improves out-of-sample R2/RMSE, correlates with an LLM-based score, and follows sensible patterns with construction era, neighborhoods, and demolition status. The concern here is not that the score lacks predictive power; it is that the paper's central interpretive claim—that the score measures 'objective product quality'—depends on the score reflecting the physical dwelling rather than the marketing presentation of the listing. The manuscript itself shows that the sample is dominated by professional agents/agencies, yet no author-type control appears in the hedonic or hazard models. Table 4 only removes technical photo-quality variation, not staged content, angle selection, or professional effort. This is a concrete, feasible-to-test confound, and it is the weakest link in the strongest claim. If the CLIP coefficient survives author-type controls and the owner-only sample, the quality interpretation is much better supported. If it does not, the paper should be reframed as measuring listing presentation, which would still be valuable but would not support the 'objective product quality' language. The reader's verdict of CONDITIONAL already captures this uncertainty, so I do not move the verdict; this stress-test provides a sharper, testable version of the same concern.","tokens_in":14758,"tokens_out":8309,"duration_ms":103330,"concrete_test":"Re-estimate Table 5 specifications (2) and (4) with a full set of author-type fixed effects (private owner, individual agent, agency account) and, separately, on the owner-only subsample (~17% of ads). If the adjusted-CLIP coefficient drops by more than roughly 30% relative to 0.582/0.350, or loses significance, the score is substantially a proxy for professional listing effort rather than physical quality. Also report the same for the Section 6 hazard model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's distinctive claim is that the CLIP Q-score measures the property's objective quality, not the way it is marketed. Section 3 (footnote 10) reports that only ~17% of ads are posted by private owners; the remaining ~83% come from individual agents or agencies. The hedonic models in Table 5 control for unit, building, geography, and district fixed effects, but not for author type, and the image-quality adjustment in Eq. (1)/Table 4 removes only technical photo characteristics (resolution, aspect ratio, sharpness, BRISQUE, brightness). Professional agents can systematically influence the visual content—staging furniture, choosing angles, lighting, and selecting which rooms to show—without changing the physical dwelling. If agencies do this more for listings they expect to price high, the CLIP score will be correlated with asking price even if physical quality is constant. The same confound affects the time-on-market hazard models in Section 6, where agent behavior (ad updating/closure) is not modeled. Thus the headline coefficients (0.582 rental, 0.350 sales per 0.1 score) and the out-of-sample R2 gain could reflect presentation/effort rather than product quality. The paper's predictive contribution would survive a reframing as 'listing presentation score,' but the title/abstract's 'objective product quality' claim would not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the CLIP Q-score, a polarity measure derived from a frozen CLIP model that compares property images against positive, neutral, and negative text prompts. The score is adjusted by residualizing on technical image characteristics (resolution, aspect ratio, sharpness, BRISQUE, brightness) and then averaged per property. The authors apply this measure to roughly half a million images from Moscow real estate listings, validate it against LLaMA4 ratings and descriptive patterns (repair type, construction-year U-shape, demolition program, district geography), and embed it in log-linear hedonic price regressions, gradient boosting models, and proportional hazard models for time on the market. They report that a 0.1 higher adjusted score is associated with roughly 5.8% higher rental prices and 3.5% higher sales prices, with modest out-of-sample R-squared gains, and that higher scores predict faster sales conditional on the asking price.","tokens_in":15073,"tokens_out":5249,"duration_ms":68657,"significance":"The methodological contribution is potentially valuable: it is fully reproducible, computationally cheap, does not export data, and is portable across markets. The empirical work is careful in several respects: train/test splits, clustered standard errors, a comparison with an LLM-based alternative, and a rich set of hedonic controls. The claim that a purely visual signal adds predictive power beyond standard observables is well supported as an association. However, the larger claim that the score measures 'objective product quality' is not yet established. The score is computed from listing images that are produced by sellers and agents, and the paper does not separate physical quality from presentation effort. That distinction is central to the paper's abstract and title, so the current version overreaches. If the score is instead interpreted as a listing-presentation or marketing-effort measure, the predictive findings remain interesting but the contribution is materially different.","major_comments":[{"comment":"The headline coefficients (0.582 rental, 0.350 sales per 0.1 score) may reflect listing presentation rather than physical quality. The adjusted score residualizes only technical image quality (resolution, aspect ratio, sharpness, BRISQUE, brightness), not the content of the photographs: staging, furniture, angle selection, lighting, and choice of rooms shown. Figure 3 and footnote 10 indicate that roughly 83% of ads are posted by professional agents or agencies, yet author type is not included in the hedonic or hazard regressions. If agents present higher-priced properties more favorably, the score is correlated with price even holding physical quality constant. The authors should add author-type controls or an owner-only subsample, or explicitly reframe the measure as a 'listing presentation score' rather than 'objective product quality'.","section":"Section 5.1, Table 5; Eq. (1) in Table 4"},{"comment":"The validation against LLaMA4 is not independent of the presentation channel: both scores are computed from the same listing images, so high correlation between CLIP and LLaMA4 is consistent with both capturing marketing effort. The descriptive evidence in Section 4.2 — repair condition, construction-year U-shape, demolition program, district geography — is more persuasive but still relies on the same images. The paper needs either an external benchmark (e.g., physical inspection, repeated listings of the same unit, or independent interior measurements) or a clear statement that the score is a presentation measure. Without this, the 'objective quality' interpretation remains under-supported.","section":"Section 4.1, Figure 7"},{"comment":"The CLIP Q-score depends on the researcher's choice of text prompts ('luxurious', 'ordinary', 'dilapidated'). The paper asserts that 'any set of text inputs are bound to yield reasonable results' but reports no sensitivity analysis over alternative prompt wording. Because the abstract claims the method extracts 'objective product quality metrics', the metric should be robust to reasonable prompt variation. Adding prompt-robustness checks, or at least a discussion of how prompt choice affects scores, would be necessary to support the method's generality.","section":"Section 2.1, Table 1"}],"minor_comments":[{"comment":"The abstract contains a typo: 'CLIP Q-store' should be 'CLIP Q-score'.","section":"Abstract"},{"comment":"'millon' should be 'million'.","section":"Section 3.1, Table 3"},{"comment":"The standard errors are labeled 'HC3' but the table does not explain the correction or how it is implemented. Also, because the adjusted CLIP score is an average residual from the first-stage regression in Eq. (1), the second-stage standard errors do not account for estimation of the first-stage parameters. Given the small first-stage R-squared this is unlikely to change the conclusions, but a bootstrap or formal correction would strengthen inference.","section":"Section 5.1, Table 5"},{"comment":"The Shapley value decomposition is useful, but the figure omits confidence intervals. Since the authors emphasize the ranking of feature importance, some uncertainty quantification would help interpret differences between adjacent features.","section":"Section 5.3, Figure 15"},{"comment":"The proportional hazard model conditions on the initial asking price. The paper explains this as holding price constant, but it would be helpful to also report the reduced-form hazard effect without price to make explicit the total effect of quality on liquidity.","section":"Section 6, Figure 17"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable in scope for an empirical economics journal. The main concern is not the quality of the predictive exercise but the interpretation of the score as 'objective quality'. The revision should either add convincing controls for the marketing channel or reframe the contribution accordingly. I would not recommend rejection because the predictive claims are competently executed and the method is genuinely reproducible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bill — This paper gives applied researchers a genuinely cheap trick: from a frozen CLIP model, three sentences (\"luxurious/ordinary/dilapidated apartment\"), and the polarity P(+)−P(−), you get an image score that survives contact with Moscow listing data. It raises rental and sale asking prices by 5.8%/3.5% per 0.1 score after hedonic controls, and it improves out-of-sample R2 by about one point. That part is real. The score is computed from a frozen model, not fitted to prices, so the worry about circularity is unwarranted.\n\nThe paper's honest strength is its validation stack: agreement with LLaMA4 scores, sensible gradients by floor/ceiling/repair, the demolition-program gap, and a U-shaped construction-year pattern that matches Moscow's history. The data collection (500k images, survival tracking) looks careful, and the econometrics is standard and competently done.\n\nThe soft spot is interpretive, not predictive. The abstract sells the score as \"objective product quality,\" but the images are mostly posted by professional agents (83%), and the regression never controls for author type. Agents stage apartments, choose angles, edit photos, and decide which rooms to show. Table 4 only removes technical photo properties (resolution, sharpness, BRISQUE, brightness). So the CLIP Q-score plausibly bundles physical condition with listing presentation. The price/liquidity predictions survive that reframing — \"listing presentation score\" is still a useful variable — but \"objective product quality\" does not. This is the one load-bearing weakness, and it's addressable: add owner/agent/agency dummies (or re-run on the 17% of owner posts), and see if the coefficient holds.\n\nTwo minor points: the paper says \"fully reproducible\" and \"all data/code open source,\" but the arXiv text has no link to code or data — that claim is currently unverifiable. And the paper cites previous work on image-based valuation (Kostic, Deng, Tapia) without any quantitative comparison to those baselines. Both are easy fixes.\n\nThis isn't a field-reorganizing paper, but it's a genuinely useful application that deserves a serious referee. I'd ask for the author-type robustness check and a more careful framing of what the score measures. Then it's publishable. Send it out.","headline":"A cheap, reproducible image score that clearly predicts Moscow rents and sale prices — but the 'objective quality' reading needs author-type controls or a reframing to 'listing presentation.'","tokens_in":15526,"tokens_out":2943,"would_cite":false,"duration_ms":34590,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces the CLIP Q-score, a fully reproducible image-based quality metric, and argues that it predicts housing prices and time on market in a large Moscow real estate dataset.","keywords":["CLIP Q-score","computer vision","multimodal machine learning","image data","hedonic price model","real estate valuation","product quality","time on market"],"falsifier":"Find the same apartment listed twice with different photo sets (professional staged photos vs. phone snapshots) and compute the CLIP Q-score for each; if the score moves systematically with photography while the physical unit is unchanged, the metric is not an objective quality measure.","tokens_in":14642,"feed_emoji":"📷","tokens_out":7035,"duration_ms":65159,"temperature":0.7,"pith_summary":"The paper introduces the CLIP Q-score, a way to turn any product photo into a quality measure by asking a pre-trained vision-language model how much closer the image is to a 'luxurious' than to a 'dilapidated' caption, with an 'ordinary' caption as a neutral anchor. Applied to roughly half a million photos of Moscow rental and sales listings, the score is argued to capture objective quality: it lines up with self-reported repair state, building age, demolition status, neighborhood prestige, and with scores from a multimodal LLM. In hedonic price regressions, a 0.1 higher adjusted score is associated with about 5.8% higher rent and 3.5% higher sale price, and including it improves out-of-sample fit. Conditional on asking price, higher scores are associated with faster sales, though not faster rentals. The paper's broader claim is that this fully reproducible, local, nearly free procedure generalizes as a general-purpose image-based quality metric for digital marketplaces.","feed_headline":"Photo-based score predicts rents, prices, and time on market","feed_subtitle":"One cheap, fully reproducible image metric adds predictive power beyond standard apartment characteristics.","key_machinery":"The load-bearing object is the CLIP polarity score computed from contrastive language-image pre-training. CLIP is a pair of image and text encoders trained so that matching image-caption pairs get nearby embeddings; the paper exploits this by feeding each photo together with three hand-written quality captions ($+$, $\\sim$, $-$), converting the three cosine similarities into probabilities via the model's learned temperature, and taking $P(+)-P(-)$ as the quality score. The same embedding machinery makes the score fully deterministic and reproducible, and the average of the per-image residuals from a regression on technical photo properties becomes the adjusted property-level score used in th","core_discovery":"On the paper's own terms, the central discovery is that the similarity structure of CLIP embeddings contains a usable quality signal without any fine-tuning. For each image, the paper computes the probability CLIP assigns to three fixed captions — 'a photo of a luxurious apartment', 'a photo of an ordinary apartment', 'a photo of a dilapidated apartment' — and defines the CLIP Q-score as the probability of the positive caption minus the probability of the negative one. Averaged over a listing's photos and residualized against technical image properties, this score is a strong predictor of log asking prices for both rentals and sales, ranked third among predictors for rentals and fourth or fi","pith_inferences":["The paper leaves open whether the score measures physical quality or the seller's presentation effort; a natural extension would compare scores for the same dwelling listed with different photo sets.","Editorial extension: the same polarity construction could be turned into a market-level index of listing glamour, letting researchers separate image-induced price effects from physical renovation effects.","Editorial extension: because the score is computed from photos, it could be used to test whether sellers strategically choose photos (e.g., omit damaged rooms), which would bias the score upward for poorly presented but structurally similar units.","If the score tracks listing presentation, its predictive power may vary across platforms with different photo norms, a testable prediction beyond the paper's Moscow sample."],"forward_implications":["Adding the CLIP Q-score to a standard hedonic model improves out-of-sample $R^2$ and RMSE for both rentals and sales, so the score is a usable low-cost feature in automated valuation.","The score works with any set of contrastive text descriptions, so the same procedure transfers to other product categories without retraining.","Because inference is local and deterministic, researchers can publish exact scores and avoid the cost, latency, and privacy issues of querying external LLM APIs.","Conditional on asking price, a higher score predicts faster sale, linking the image measure to liquidity in search-and-matching models of housing.","If the results hold, asking prices understate the score's relationship to transaction prices, since better-looking units give owners bargaining power."],"supporting_citations":[{"why":"Supplies the pre-trained CLIP model and the contrastive embedding geometry that the polarity score is built on.","marker":"Radford et al., 2021"},{"why":"Formalizes using CLIP embedding similarity as a reference-free quality score, the direct precursor of the paper's polarity measure.","marker":"Hessel et al., 2021"},{"why":"Defines hedonic prices as implicit attribute prices, the theoretical framework for the price regressions.","marker":"Rosen, 1974"},{"why":"Establishes the precedent of enriching hedonic real estate models with non-traditional textual data, which the paper extends to image data.","marker":"Nowak and Smith, 2017"},{"why":"Shows image-derived features boost housing price predictions, giving the paper a benchmark for image-based valuation.","marker":"Kostic and Jevremovic, 2020"},{"why":"Demonstrates that product images predict return rates, supporting the claim that images carry economically relevant quality information.","marker":"Dzyabura et al., 2023"},{"why":"Documents non-determinism in LLM outputs, motivating the paper's reproducibility advantage of CLIP over LLM quality scores.","marker":"Atıl et al., 2025"},{"why":"Provides institutional detail on Moscow's mass-housing demolition program used as a validation pattern for the score.","marker":"Gunko et al., 2018"}],"fun_headline_variants":["CLIP Q-score from images predicts housing prices and rentals","No fine-tuning needed: CLIP embeddings forecast real estate values","Image quality score via CLIP beats standard home features","Three captions, one score: CLIP predicts time on market","Frozen CLIP model yields objective product quality metric"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The score is assumed to reflect the apartment's actual physical quality rather than how the listing was photographed, staged, or marketed; if high scores mostly capture presentation, the price and duration coefficients describe marketing, not the dwelling.","fun_headline_variants_meta":{"raw":{"variants":["CLIP Q-score from images predicts housing prices and rentals","No fine-tuning needed: CLIP embeddings forecast real estate values","Image quality score via CLIP beats standard home features","Three captions, one score: CLIP predicts time on market","Frozen CLIP model yields objective product quality metric"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":996,"prompt_tokens":609,"completion_tokens":387,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":353,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":353,"tokens_out":387,"duration_ms":5009,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:03:29.073068+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find the same apartment listed twice with different photo sets (professional staged photos vs. phone snapshots) and compute the CLIP Q-score for each; if the score moves systematically with photography while the physical unit is unchanged, the metric is not an objective quality measure.","supporting_citations":[],"review_version":1}