{"id":"7d184e5a-4143-4fc0-a756-f3429cef6e74","arxiv_id":"2504.17113","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A nine-bedroom coliving house sustained 98% occupancy and below-market rents for 18 months using a points-based chore auction, a peer 'hearts' ledger, and shared purchasing, all without a manager.","lead":"A Los Angeles coliving house ran for 18 months with no manager, using custom digital tools to assign chores, track peer appreciation, and approve shared purchases. The case study offers a concrete template for how small self-governing groups can use simple feedback systems instead of leaders.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sage House's success is not identifiable as an effect of Chore Wheel, because the owner's residual role, the monthly in-person meeting, and resident selection are confounds the data do not separate.","rationale":"The reader's weakest assumption identifies exactly the concern that matters: the absence of a causal contrast separating the Chore Wheel mechanisms from the owner's residual role, the monthly in-person meeting, and the particular resident culture. The paper is honest about this in Section 5, which weakens the claim of 'validation' in the same section but does not make the paper an internal contradiction. The mechanisms are described concretely, the data are presented as exploratory, and the open-source implementation is checkable, so the paper retains value as a design case study. However, the abstract's and Section 5's stronger claim that such tooling 'can sustain reliable commons governance without continuous leadership' is not supported by the present evidence. The verdict should remain conditional: accept the descriptive contribution and design principles, but withhold the causal claim pending a comparison that separates tool effects from human backstops. My proposed test is feasible because the owner lived in the house for the first year, providing a natural within-case contrast that the paper does not report.","tokens_in":8963,"tokens_out":3855,"duration_ms":39779,"concrete_test":"Re-analyze the existing Slack and Chore Wheel logs to compare the first 12 months, during which the owner lived in the house, with the subsequent months after the owner moved out. Report monthly occupancy, chore completion rate, hearts lost to shirking, and the number of owner-initiated administrative messages (account changes, dispute resolutions, manual point edits). If occupancy or chore compliance drops significantly after the owner's departure, or if owner-admin messages continue at a nontrivial rate under the 'limited administrative role,' the tool-alone attribution fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the Chore Wheel mechanisms sustained reliable commons governance without continuous leadership. The load-bearing problem is that the 18-month outcome cannot be causally assigned to the tooling, because the case includes three unseparated human backstops. Section 1 states that the owner co-owns the house, lived there for its first year, and 'continues to play a limited administrative role'; Section 3.3 describes a 'constitutional action' that added named accounts; and Section 5 concedes that residents hold a monthly in-person meeting whose role is 'likely essential' to handling residual issues, while also asking directly, 'What is the role of abstract structure, and what is the role of specific culture, in explaining Sage's performance?' The exploratory data (567 chore claims, 255 heart events, 551 purchases) show activity but not counterfactual dependence: there is no baseline, no control house, and no comparison of periods with and without the owner in residence. If the owner's residual administration, the monthly meetings, or the particular conscientiousness of the 16 selected residents did the coordinating work, then the headline claim that 'such tooling can sustain reliable commons governance without continuous leadership' is not established. The paper's own caveats make this a missing-evidence concern rather than an internal contradiction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an 18-month field observation of a nine-bedroom Los Angeles coliving house ('Sage House') governed by a Slack-based software suite called Chore Wheel. The authors describe three mechanisms: a continuous-auction chore scheduler with a 100-point monthly obligation, a pairwise-preference layer for reprioritizing chores, and a symbolic 'hearts' ledger for norm enforcement. They frame the design using Ostrom's commons principles and cybernetic theory, and they present exploratory descriptive statistics from the house: 567 chore claims, 255 heart events, and 551 group purchases. The paper's central claim, as stated in the abstract and discussion, is that this tooling can sustain reliable commons governance without continuous leadership. The authors explicitly acknowledge in Section 5 that the case has confounds and that the analysis is exploratory, including the roles of the owner, resident composition, and a monthly in-person meeting.","tokens_in":9233,"tokens_out":3358,"duration_ms":32753,"significance":"If the central claim were established, the paper would be a valuable real-world existence proof that computational institutional mechanisms can reduce the leadership burden in a housing commons, with implications for Ostrom-inspired institutional design, online communities, and coliving. The paper has several genuine strengths: the mechanisms are concretely specified, the software is open-source, the data are anonymized and made available, and the authors are unusually candid about limitations and confounds. However, as presented, the paper is best understood as a design case study and feasibility demonstration rather than a causal validation of the tooling, because the observational design cannot separate the effect of Chore Wheel from the effect of the specific residents, the owner's residual administrative role, and the supplemental in-person meeting. Its main value lies in articulating a transferable design palette and generating hypotheses, not in proving that the tooling alone sustains governance.","major_comments":[{"comment":"The abstract and Section 5 claim that the house 'runs without managers' and that the tooling can sustain governance 'without continuous leadership,' but the manuscript itself documents three continuous human backstops: Section 1 states that the owner co-owns the house, lived there for its first year, and 'continues to play a limited administrative role'; Section 3.3 describes a 'constitutional action' that introduced named accounts; and Section 5 concedes that the monthly in-person meeting 'plays what is likely an essential complementary role.' These are precisely the kinds of ongoing human coordinating functions that the paper needs to rule out in order to attribute Sage's performance to Chore Wheel. Because no period without these backstops is analyzed, the headline causal claim is not supported by the reported evidence. Please either provide comparative evidence (e.g., periods with and without the owner in residence, or without the monthly meeting) or reframe the contribution as a design feasibility study rather than a demonstration that the tooling alone sustains governance.","section":"§1 and §5"},{"comment":"The exploratory data (567 chore claims, 255 heart events, 551 purchases) document that the system was used, but they do not by themselves show that the house 'functioned as well or better than pure-human alternatives' (Section 5) or that the outcomes are attributable to the mechanisms. There is no baseline period, no control house, and no comparison of outcomes with and without Chore Wheel; occupancy rate, below-market rent, and tenure are reported without a counterfactual. The paper acknowledges 'many possible confounds' in Section 5, but the same section concludes that the case 'represents a validation of Chore Wheel's theory and practice.' That conclusion outruns the evidence. I suggest replacing the validation language with a more measured statement that the case is consistent with the proposed design theory and provides a proof-of-concept, and explicitly listing the conditions under which the design would be falsified.","section":"§4 and §5"},{"comment":"The core modeling assumption of the Chores mechanism—that a continuous, linear increase in points over time 'effectively models stochastic, non-linear increases in mess and disorder'—is stated without empirical support, yet it is load-bearing for the claim that the continuous-auction scheduler works. The paper also does not discuss how the linear accrual rate is calibrated or how sensitive the mechanism is to the choice of rates, the fixed 100-point obligation, the baseline of five hearts, or the adaptive approval thresholds. Since the authors claim transferability, an explicit discussion of these parameters and at least a sensitivity check or robustness argument would materially strengthen the paper. If the assumption is intended only as a design heuristic, it should be labeled as such rather than presented as a modeling result.","section":"§3.1"}],"minor_comments":[{"comment":"The abstract says '18-month field experiment,' while Section 5 says the study is 'based on 20 months of field data.' Please reconcile these numbers.","section":"Abstract and §5"},{"comment":"There are typographical errors, including 'Succintly' in Section 1, 'go to mangement' in Section 1, and a stray '-blur' at the end of Section 3.3. These should be corrected.","section":"Throughout"},{"comment":"The normalization of monthly chore performances is described only briefly in the caption. Please specify the normalization formula and state what the y-axis represents, since the reader cannot otherwise interpret the comparison across residents with different total points.","section":"Figure 3"},{"comment":"The term 'naturally affordable' is used to describe Sage's rents, but it is not defined. Please provide a precise definition or reference, especially since housing affordability is mentioned as a significant implication.","section":"§1"},{"comment":"The analogy to 5,000 Minecraft servers (Frey and Sumner, 2019) is intended to support a directional claim, but the leap from large-scale server data to this single-house case is not self-evident. Please add one or two sentences explaining how that result bears on the present case.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a self-reported case study in which one author is the designer and co-owner of the system and the house. This is disclosed in the ethics statement, but the overlap between system developer, operator, and ethnographer is a substantial reporting-bias risk that the public review alone may not fully convey. I would recommend asking the authors to make the anonymized raw data and analysis scripts available for review and to state explicitly what independent verification, if any, was performed. The paper fits the scope of a journal like the one under consideration if it is presented as an exploratory design case study; the current overclaim in the abstract and conclusion is what pushes it to major revision rather than acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a case study of Sage House, a nine-bedroom LA coliving house run with a Slack-based \"Chore Wheel\" for chores, hearts, and purchasing. The real contribution is the 18-month descriptive dataset: 567 chore claims, 255 heart events, 551 purchase proposals, with 98% occupancy and below-market rents. That dataset is new, and the paper is refreshingly honest about its limits—Section 5 explicitly asks whether abstract structure or resident culture explains the performance, and concedes the monthly in-person meeting plays a role it calls \"likely essential.\"\n\nWhat it does well: design principles are clearly laid out, the mechanisms are described concretely enough to reproduce (software is open source), and the exploratory analyses—task specialization, hearts dynamics, purchase power laws—give the field something to chew on. The paper also does not hide the owner's dual role or the absence of a control group.\n\nWhere it's soft: the central claim in the abstract—\"such tooling can sustain reliable commons governance without continuous leadership\"—is not actually established by the evidence. The house had three human backstops that are not separated from the tooling: the owner (a co-author) lived there for the first year and continues a limited administrative role; residents hold a monthly in-person meeting that the paper calls essential; and the 16 residents were self-selected, likely conscientious people. There is no baseline, no comparison house, no period without the owner. The paper's own Minecraft-servers analogy (Frey and Sumner) cuts the other way: that argument relies on a large-N correlation, while this is n=1.\n\nThe exploratory data are fine as a proof-of-possibility, and the authors deserve credit for flagging the confounds rather than burying them. But the abstract's \"show that\" overstates what a single case with this design can show. I'd reframe it as \"suggests\" or \"is consistent with.\"\n\nWho's it for: people working on coliving, DAO tooling, and Ostrom-inspired digital governance. They'll get a concrete, candid field example and a few testable hypotheses. It deserves a serious referee—with revisions to moderate the causal language and perhaps add a preregistered replication plan. Worth a careful read, worth citing for the dataset and the limitations discussion.\n\nRecommendation: send it to peer review, but expect and require toning down of the causal claim.","headline":"A useful, honest case study with a new descriptive dataset, but the abstract's causal claim that the tooling sustains governance without leadership is not actually supported by the evidence.","tokens_in":9754,"tokens_out":1676,"would_cite":true,"duration_ms":15893,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A nine-bedroom Los Angeles coliving house sustained 98% occupancy for 18 months with no manager, using a digital chore-auction, hearts-ledger governance toolkit.","keywords":["coliving","commons governance","cybernetic governance","chore allocation","continuous auction","pairwise preference","hearts ledger","self-governance"],"falsifier":"Compare two nearly identical coliving houses over 18 months, one using the Chore Wheel continuous-auction scheduler and the other using a fixed chore rotation with the same chores and no digital mechanism; if the fixed-rotation house matches 98% occupancy, cleanliness, and below-market rents, the specific auction and ledger mechanisms are not the operative cause.","tokens_in":8736,"feed_emoji":"🏠","tokens_out":8495,"duration_ms":74450,"temperature":0.7,"pith_summary":"This paper reports an 18-month field experiment in a nine-bedroom Los Angeles coliving house that ran without managers while sustaining 98% occupancy and below-market rents. The authors argue that a digital toolkit called Chore Wheel replaced much of the coordinating work a manager would do: chores became a time-indexed points market, residents reprioritized tasks through simple pairwise comparisons, and a symbolic hearts ledger tracked norms. The mechanisms are tied to established design principles for governing shared resources, and exploratory data from the house cover 567 chore claims, 255 heart events, and 551 group purchases. If the claim holds, coliving houses, online communities, and other digitally mediated collectives have a transferable design palette for sustaining commons governance without continuous leadership.","feed_headline":"98% occupancy, no manager: 18 months of cybernetic house rules","feed_subtitle":"A digital chore auction, hearts ledger, and group purchasing kept a nine-bedroom coliving commons running without a leader.","key_machinery":"The central object is the Chore Wheel suite, implemented as apps in a shared communications platform. Its first mechanism is a continuous-auction chore scheduler: every recurring task accrues points linearly with time since it was last done, residents owe 100 points per month, and claims are verified by community upvotes and downvotes. The second is a pairwise-preference layer that lets residents asynchronously state which tasks deserve higher priority, generating dynamic priority distributions that control how quickly each chore gains points. The third is a symbolic hearts ledger that tracks norm compliance through peer-awarded karma, lightweight challenges, automatic penalties, and passive regeneration over time. A simpler purchasing mechanism, Things, uses price-scaled approval thresholds for shared spending. Together these mechanisms perform the bookkeeping and aggregation while residents contribute low-cognition judgments from local information.","core_discovery":"The central claim is that, in one real-world housing commons, a suite of computationally mediated mechanisms sustained reliable governance over 18 months without a manager or privileged roles. The mechanisms operationalize a cybernetic loop in the paper's sense: humans sense and report local conditions, machines keep the books and compute priorities, and feedback adjusts incentives in real time. The paper does not claim leadership disappeared, only that tasks historically requiring personal leadership—scheduling regenerative labor, setting priorities, enforcing norms, and managing shared purchases—could be accomplished by non-leaders using the system. The finding is presented as exploratory, with the house's continuous occupancy at below-market rents offered as evidence that the governance held.","pith_inferences":["If the mechanism is causal, then adding neglect-based point accrual to existing community task tools could reduce coordinator burnout in volunteer groups, with no custom software needed for houses that already use shared chat platforms.","The strongest test would be a matched-house experiment or an internal switch: run the continuous-auction chore scheduler against fixed chore values for several months and measure task delay, cleanliness, and resident satisfaction.","The named-accounts episode suggests that institutional outcomes can be changed by relabeling the very same mechanism, which points to a general design lever: changing the semantic context of a decision can be as powerful as changing incentives.","The authors acknowledge that Sage House's residents were unusually capable, so whether the toolkit sustains a less selective population is the open empirical question this study cannot answer."],"forward_implications":["A nine-bedroom coliving house sustained 98% occupancy and below-market rents over 18 months with no manager, and residents met a monthly obligation of roughly 2-4 hours of chores plus one 90-minute meeting.","The mechanisms let residents self-specialize in tasks without explicit organization, and the hearts ledger appeared to channel interpersonal tension into a regenerative symbolic system rather than stored grievances.","The design principles generalize beyond housing: the paper explicitly offers them as a transferable palette for online communities and other digitally mediated collectives.","Leadership is reduced and distributed rather than eliminated; the paper reports that a monthly in-person meeting remains an essential complement to the technical system.","Mechanistically identical purchasing accounts produced savings behavior only after they were given names, indicating that symbolic semantics shape how residents use the same technical machinery."],"supporting_citations":[{"why":"Supplies the eight design principles for governing shared resources that Sage's mechanisms realize.","marker":"[14]"},{"why":"Provides the institutional-analysis framework the paper draws on for its four additional design principles for distributed digital institutions.","marker":"[15]"},{"why":"Introduced the Chore Wheel suite and anchors the design lineage of the three mechanisms.","marker":"[9]"},{"why":"Gives the pairwise-preference analysis used to turn asynchronous individual comparisons into task priorities.","marker":"[7]"},{"why":"Extends pairwise preferences to a resource-allocation mechanism, the basis of the distributed prioritization layer.","marker":"[8]"},{"why":"Establishes the static fair-chore-division baseline that the continuous auction generalizes to dynamic settings.","marker":"[1]"},{"why":"Provides an iterated fair-allocation result that the paper moves beyond to handle drifting task valuations.","marker":"[5]"},{"why":"Offers comparative evidence from many self-governing communities that integrated institutions correlate with scale, supporting the paper's direction of causality.","marker":"[2]"}],"fun_headline_variants":["No manager, 98% full: coliving runs on digital rules","Cybernetic coliving: 18 months without a leader","Chore auction and hearts ledger sustain manager-free coliving","Digital commons keep coliving house 98% occupied","Ostrom's principles, digital tools: coliving self-governs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The house's stable performance is attributed to the Chore Wheel mechanisms, but the paper leaves open that the same outcome could have come from the particular, unusually capable residents or from the owner's continued administrative oversight.","fun_headline_variants_meta":{"raw":{"variants":["No manager, 98% full: coliving runs on digital rules","Cybernetic coliving: 18 months without a leader","Chore auction and hearts ledger sustain manager-free coliving","Digital commons keep coliving house 98% occupied","Ostrom's principles, digital tools: coliving self-governs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1382,"prompt_tokens":896,"completion_tokens":486,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":512,"tokens_out":486,"duration_ms":4414,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:48:44.327971+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare two nearly identical coliving houses over 18 months, one using the Chore Wheel continuous-auction scheduler and the other using a fixed chore rotation with the same chores and no digital mechanism; if the fixed-rotation house matches 98% occupancy, cleanliness, and below-market rents, the specific auction and ledger mechanisms are not the operative cause.","supporting_citations":[{"cited_title":"Governing the Commons","cited_arxiv_id":null,"evidence_quote":"Supplies the eight design principles for governing shared resources that Sage's mechanisms realize."},{"cited_title":"The new chicago school","cited_arxiv_id":null,"evidence_quote":"Provides the institutional-analysis framework the paper draws on for its four additional design principles for distributed digital institutions."},{"cited_title":"Mirror whitepaper,","cited_arxiv_id":null,"evidence_quote":"Introduced the Chore Wheel suite and anchors the design lineage of the three mechanisms."},{"cited_title":"An Analysis of Pairwise Preference","cited_arxiv_id":null,"evidence_quote":"Gives the pairwise-preference analysis used to turn asynchronous individual comparisons into task priorities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends pairwise preferences to a resource-allocation mechanism, the basis of the distributed prioritization layer."},{"cited_title":"How to fairly allocate easy and difficult chores","cited_arxiv_id":null,"evidence_quote":"Establishes the static fair-chore-division baseline that the continuous auction generalizes to dynamic settings."},{"cited_title":"Repeated fair allocation of indivisible items","cited_arxiv_id":null,"evidence_quote":"Provides an iterated fair-allocation result that the paper moves beyond to handle drifting task valuations."}],"review_version":1}