Pith. sign in

REVIEW 3 major objections 5 minor 89 references

SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that SocialFiVis lets operators test counterfactual governance policies and trace macro-level shifts to individual LLM-grounded personas.

desk verdict A competent, well-integrated VA sandbox whose headline emergent finding rests on the one trust metric that fails validation in the exact community used to show it; worth peer review, but the emergent claims need re-scoping or an ablation. read the letter →

arxiv 2608.08497 v1 pith:4D6CMYOC submitted 2026-08-09 cs.HC cs.MAcs.SI

classification cs.HCcs.MAcs.SI
keywords SocialFivisualanalyticsLLM-groundedagentsimulationcounterfactualreasoningdigitalcommonsgovernancemulti-agentcapitalIADframework
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Social finance communities blend social interaction with token economies, so governing them is high-stakes: real interventions cost money and can irreversibly damage trust or liquidity. This paper builds a visual analytics sandbox, SocialFiVis, that lets community operators inject hypothetical governance policies into a simulated community and watch macro-level metrics respond. The paper's central claim is that the sandbox supports fine-grained behavioral attribution: operators can trace shifts in social capital and financial health to the rationales of specific simulated personas. In case studies the system surfaces emergent patterns, including a structural decoupling in which trust drops while participation and consensus hold, and a resilience of messaging members to localized governance shocks. A sympathetic reader would care because the system turns irreversible real-world experiments into risk-free counterfactual backtests with a visible chain from policy to individual cognition to collective outcome.

What carries the argument

The load-bearing machinery is the two-phase simulation engine paired with a closed-loop metric feedback. Phase I extracts personas by clustering retained messaging users with K-Modes over a seven-dimensional trait codebook; Phase II runs a mechanism-guided Perception–Reasoning–Action (PRA) pipeline in which agents consult a five-layer memory stack and act asynchronously under a coordinator that regulates turn-taking. Simulated actions feed back tick-by-tick into the quantitative definitions of the commons: social capital uses a soft-penalty geometric blend $\mathrm{SC}'_t = \alpha\cdot\text{arith} + (1-\alpha)\cdot\text{geom}$ with $\alpha=0.5$, and financial health uses an unweighted geometric mean $\mathrm{FH}_t = \sqrt[3]{F_t H_t L_t}$. This closed loop is what lets an intervention propagate from an individual agent's reasoning to macro-level metric shifts, and what lets the interface trace the shifts back to personas.

What would settle it

Re-run the two case studies with the ground-truth-anchored environmental inputs withheld, so that activity levels, sentiment, and topic concentration are not fed from historical data, and check whether the trust–participation decoupling and the pessimistic-persona sell-off survive. If they vanish, they are calibration artifacts rather than emergent behavior. A complementary test is to apply the pipeline to a real governance change that occurred after the study window and compare simulated trajectories with the observed metric shifts.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that an institutional framework can be operationalized as a closed-loop, LLM-grounded simulation pipeline that makes emergent socio-financial phenomena attributable. The system quantifies a dual-track digital commons—social capital from participation, consensus, and trust, and financial health from floor price, liquidity, and holder count—then instantiates heterogeneous personas from a seven-dimensional codebook and runs them through a Perception–Reasoning–Action runtime under user-injected governance rules. The reported case studies show the pipeline revealing diminishing returns from stacked incentive policies, covert exploitation by personas whose stated sentiment diverges from their trades, and an isolated trust decline under a localized negative shock that aggregate engagement would mask. The paper presents these findings as explanatory, attribution-supporting outcomes rather than forecasts, and grounds them by validating the no-intervention mode against historical ground truth.

Load-bearing premise

The simulation is assumed to stay informative about the real community under new policies because its no-intervention mode tracks historical ground truth; if that match largely reproduces calibrated inputs, the reported emergent phenomena could be artifacts of the anchoring.

Editorial extensions

If this is right

  • Community operators can compare counterfactual governance policies against the historical baseline without spending real budgets, converting strategy intuition into testable backtests.
  • Aggregate metric movements become attributable: a drop in trust can be inspected down to the personas that sold and refuted peers, rather than remaining an anonymous aggregate shift.
  • The diminishing-returns result implies that stacking incentive policies can dilute consensus, so staggering incentive releases is a directly actionable policy design rule.
  • The sentiment–action divergence detected in a persona suggests monitoring for manipulative archetypes during incentive campaigns, since stated optimism can accompany aggressive selling.
  • The authors themselves bound the claim: outcomes are exploratory reasoning aids, not forecasts, and the simulation reflects the retained messaging cohort, not silent members.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the no-intervention fidelity transfers to counterfactual validity, the same institutional pipeline should generalize to other common-pool-resource communities—open-source projects, DAOs, creator economies—by swapping data streams and persona codebooks.
  • Editorial inference: the trust–participation decoupling could become a real-time early-warning diagnostic, monitored continuously rather than only in counterfactual mode.
  • Editorial inference: a decisive test the paper does not run is a post-hoc backtest against a real governance change that occurred after the study window; agreement there would materially strengthen the counterfactual case.
  • Editorial inference: adding persistent belief states separable from expression, which the paper lists as future work, would make word-action discrepancies a systematic, auditable signal rather than an incidental finding.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript presents SocialFiVis, a visual analytics sandbox for exploring counterfactual governance policies in SocialFi (NFT) communities. It operationalizes Ostrom's IAD framework into a dual-track model of social capital (participation, consensus, trust) and financial health (floor price, liquidity, holders), and combines LLM-derived personas with a mechanism-guided Perception–Reasoning–Action runtime to simulate heterogeneous agents. The system is evaluated through two expert case studies, a 13-participant user study with Likert ratings, and follow-up interviews. The paper claims that SocialFiVis supports fine-grained behavioral attribution and explains emergent phenomena such as the structural decoupling of social capital and the resilience of messaging members under localized governance shocks.

Significance. If the underlying simulation is trustworthy, this is a strong and timely contribution to visual analytics for social-financial systems: it gives community operators a risk-free backtesting environment, it explicitly couples macro economic and meso/micro behavioral views, and it is unusually transparent about its limitations, including the exclusion of silent members, the lack of persistent belief states, and the general caveat that historical fit does not guarantee predictive validity. The user study is reasonably structured, the sensitivity analyses for the metric model are a welcome addition, and the paper ships supplemental materials on OSF. The visual design choices (capsule-and-ribbon Behavior View, multi-ring Communication Network) are thoughtfully justified. However, the paper's central empirical claim about emergent trust decoupling rests on a trust metric that fails Mode0 validation in the very community used for that claim, and the counterfactual interpretation is confounded by the GT-anchored calibration design. These issues are load-bearing and need to be addressed before the central claims can be accepted.

major comments (3)
  1. [§8.1.2, Table 1, §6.2.2] Case II's headline 'structural decoupling' insight rests on exactly the metric and community for which Mode0 validation fails. Table 1 reports T_t for Mfers with ρ=0.421 and p=0.073, the only non-significant entry in the table, and §6.2.2 explicitly states 'This limits trust-specific interpretation for Mfers.' Yet §8.1.2's Actionable Insight 1 ('Watch trust–participation decoupling despite stable engagement') is built on a 'marked and isolated decline' in that same T_t in that same community, and the abstract elevates 'structural decoupling of social capital' to a demonstrated emergent phenomenon. Because the central claim depends on this instance, the authors should either provide additional validation that the simulated trust dip is reliable despite the failed Mode0 result, or reclassify this insight as an unvalidated hypothesis rather than a demonstrated finding.
  2. [§6.2.2, §8.1.2, §9.2] The counterfactual response interpreted as emergent in Case II is confounded by the GT-anchored calibration design. Section 6.2.2 states that empirically observed activity levels, sentiment, and topic concentration are fed into the simulation as exogenous environmental inputs; Eq. (1) (P_t) and Eq. (2) (C_t) depend directly on message volume, topic entropy, and sentiment variance—exactly the anchored channels—while Eq. (3) (T_t) is the channel freest to respond. A shock that leaves the anchored channels stable while the trust channel dips is therefore the default output of this architecture rather than surprising evidence of agent-level response to the injected policy. The paper provides no ablation, placebo run, or null-policy control to show that the Case II decoupling is driven by simulated agent reactions rather than by the calibration inputs. I request such a control (for example, injecting a semantically inert event, or running the same shock with the GT-anchored channels frozen) and a correspondingly cautious wording in §8.1.2. The general caveat in §9.2 that historical fit does not guarantee counterfactual validity is not sufficient, because the issue here is an internal design confound, not only the usual extrapolation risk.
  3. [Table 1, §6.2.2] The reported Mode0 validation is statistically under-specified and partly circular. FHt achieves 1.000 correlation and 0.000 JSD by design, since it uses unmodified real-world financial data, so the composite fidelity scores in Table 1 overstate the amount of independent validation. The Spearman correlations are computed on daily time series with strong autocorrelation, so the reported p-values (including the p=0.073 for Mfers T_t) are not valid evidence about trend fidelity, and no confidence intervals, number of time points, or DTW/JSD significance thresholds are reported. Since Table 1 is the paper's only quantitative support for the claim that the simulation tracks ground truth, I ask for an autocorrelation-aware test or block bootstrap, exact per-community sample sizes, and a clearer separation of validated metrics from metrics that are calibrated or fixed by construction.
minor comments (5)
  1. [§5.2, Eq. (4)] The description of α=0.5 as a 'symmetric blend' is potentially confusing; the formula is a convex combination of arithmetic and geometric means, not a symmetric operation in any usual mathematical sense. Consider calling it 'balanced' or spell out the intended symmetry.
  2. [Fig. 1] Figure 1 is extremely dense and contains many unlabeled or barely legible components (for example, the C0–C5 and L1–L5 labels). Please enlarge the figure and add a short legend or caption explanation for the main acronyms, since this figure is the primary overview of the system.
  3. [§6.2.2] The sentence 'This limits trust-specific interpretation for Mfers' is a strong and honest limitation, but it appears only after the validation table. Consider restating this caveat in the abstract or introduction so that readers do not encounter the trust-based 'structural decoupling' claim in the abstract before they see the validation result.
  4. [§9.2] The limitation paragraph correctly notes that the simulation excludes silent and near-silent accounts, and that the system 'therefore reflects expressed dynamics, not silent disengagement.' This boundary should be stated earlier, ideally in §6.1 where the retained cohort is introduced, because it directly constrains the scope of the case-study claims.
  5. [§8.2.1] The participant numbering is slightly confusing: E1 and E5 bypass the predefined tasks because of their case studies, but the reader must infer that E5 was recruited later than E1–E4. Please clarify the participant timeline in one sentence.

Circularity Check

3 steps flagged · score 6.0 of 10

Overall fidelity and the headline 'structural decoupling' rest on by-construction components: FHt is ground truth by design, P_t/C_t inherit GT-anchored inputs, and the trust channel used for the decoupling insight fails Mode0 validation in Mfers.

  1. fitted input called prediction [Sec. 6.2.2, Table 1 note; Eq. (6) FHt definition]
    "FHt matches ground truth by design in Mode0. ... FHt achieves perfect calibration in Mode0 by design, as it uses unmodified real-world financial data, establishing a principled baseline against which governance interventions are evaluated."

    The 'Overall Fidelity' composite in Table 1 includes FHt, whose JSD=0.000 and Spearman rho are perfect because FHt is literally the unmodified real-world financial data used for validation. Including a by-construction-perfect component in the composite inflates the reported overall fidelity (0.746 and 0.759) and makes 'Mode0 tracks ground truth' partly tautological. The paper discloses this, and a calibrated baseline is legitimate, but the composite fidelity score cannot be read as evidence that the simulation predicts financial health.

  2. fitted input called prediction [Sec. 6.2.2 (GT-anchored calibration) with Eqs. (1)-(3)]
    "The final version introduces GT-anchored environmental calibration, a standard ABM practice [23,74] wherein empirically observed activity levels, sentiment, and topic concentration serve as exogenous environmental inputs."

    P_t (Eq. 1) is built from daily message counts and distinct senders; C_t (Eq. 2) is built from topic entropy and sentiment standard deviation. These are exactly the channels that GT-anchored calibration feeds in as exogenous empirical inputs. Therefore the high Mode0 Spearman values for P_t (0.910) and C_t (0.758) largely reproduce the calibration inputs rather than measuring emergent agent behavior. Only T_t (Eq. 3) is free enough to be non-tautological, and it is the one metric that fails in Mfers (rho=0.421, p=0.073). The validation claim of 'strong trend fidelity for core metrics' is partially self-confirming.

1 more flagged steps
  1. fitted input called prediction [Sec. 8.1.2, Case II, Actionable Insight 1]
    "the participation and consensus metrics remained stable in the Event Timeline, maintaining levels comparable to the positive intervention. In contrast, the trust index exhibited a marked and isolated decline. This structural decoupling shows E5 that strong macro-positive signals can sustain engagement and consensus even when interpersonal trust erodes under a localized negative shock."

    The abstract's headline 'structural decoupling' is instantiated by Case II: participation/consensus stable, trust isolated decline. But the stability of P_t/C_t is inherited from GT-anchored inputs, while T_t is the unanchored channel that Table 1 shows failing validation in Mfers (p=0.073) and Sec. 6.2.2 says 'limits trust-specific interpretation for Mfers.' The paper then builds the decoupling insight directly on that trust-specific interpretation in that community. The contrast is thus the expected output of an anchoring scheme that pins two channels to ground truth and leaves the third free; no ablation or placebo run shows the trust dip is caused by agent responses to the injected policy rather than by the calibration asymmetry.

full rationale

This is not a fully circular paper: the IAD operationalization, the persona extraction and codebook validation, the sensitivity analyses, and the user study are independent contributions, and the self-citations to NFTracer/NFTeller are not load-bearing. No uniqueness theorem is imported from the authors, and the soft-penalty geometric blend is justified by AM-GM reasoning and sensitivity checks rather than by citation. The circularity is concentrated in the validation-and-insight chain. First, FHt is admitted to match ground truth by design because it uses unmodified real-world financial data, yet it is incorporated into the reported overall fidelity score. Second, GT-anchored environmental calibration feeds in exactly the activity, sentiment, and topic-concentration channels from which P_t and C_t are computed, so their strong Mode0 correlations partially measure the calibration inputs. Third, the one metric that is not anchored, T_t, fails Mode0 validation in Mfers, and the paper even states this 'limits trust-specific interpretation for Mfers'; nevertheless, Case II builds the abstract's 'structural decoupling' claim on an isolated trust decline in exactly that community. The paper is transparent about these limitations and explicitly scopes the counterfactual outcomes as exploratory reasoning aids rather than forecasts, which prevents a higher score. Even so, the headline fidelity numbers and the central emergent phenomenon reduce in part to the calibration design rather than to independent agent behavior.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on several hand-assigned parameter values and domain assumptions rather than a parameter-free derivation. The metric equations use equal weighting and a fixed blend coefficient, all justified by sensitivity analysis but not derived. The simulation's validity depends on the IAD mapping, the seven-dimensional persona codebook, LLM label reliability, and the retained messaging cohort. The most consequential assumption is that GT-anchored calibration does not compromise counterfactual inference. No genuinely new physical or formal entities are postulated; the dual-track metrics are measurement constructs rather than new ontological commitments.

free parameters (6)
  • alpha in soft-penalty blend = 0.5
    Chosen as a symmetric blend between arithmetic and geometric means in Eq. 4; sensitivity analysis shows rank order is stable, but the value is a design choice, not derived.
  • sub-metric weights in SCt (P, C, T) = equal 0.5 weights
    Eqs. 1-3 use equal splits with a 0.5 centering offset; sensitivity analysis checks weights in [0.25,0.75], but the nominal values are hand-assigned.
  • implicit reciprocity window = 5 minutes
    Defines sequential topically related posts as reciprocated edges in Sec. 5.2.1; no principled derivation is given for this time scale.
  • persona inference threshold = 15 messages
    Accounts below this threshold are excluded in Sec. 6.1; the threshold is chosen to ensure reliable LLM labeling, not derived from data.
  • number of persona clusters K = 6 for Mfers, 7 for Mimic Shhans
    Selected via inertia, silhouette, and ARI in Sec. 6.1; a model-selection choice rather than a parameter of the central derivation.
  • minimum agents per archetype = Np >= 5
    Chosen for statistical stability in Phase II population allocation, as described in Sec. 6.2.1.
assumptions (6)
  • domain assumption The IAD framework is an appropriate model for SocialFi community governance.
    Sec. 3.2 adopts Ostrom's framework based on practitioner input; if SocialFi interactions do not follow the IAD rule-action-outcome structure, the architecture loses its theoretical grounding.
  • domain assumption The seven 3-class persona dimensions capture behaviorally relevant heterogeneity.
    The codebook in Sec. 6.1 is expert-validated (Appendix Fig. 7), but the dimensions are not derived from a formal theory of SocialFi behavior.
  • domain assumption LLM persona labels are reliable enough for clustering.
    Sec. 6.1 relies on LLM classification with confidence-guided re-annotation; no inter-annotator agreement statistic is reported.
  • domain assumption The retained messaging cohort (at least 15 messages per user) is sufficient to represent community dynamics.
    Sec. 9.2 acknowledges silent members are excluded, so all simulation and case-study conclusions are bounded to expressed behavior.
  • domain assumption GT-anchored environmental calibration preserves counterfactual validity.
    Sec. 6.2.2 feeds observed activity, sentiment, and topic concentration into the no-intervention mode; the paper assumes this does not invalidate comparisons under injected policies.
  • standard math Standard mathematical tools (AM-GM inequality, Shannon entropy, LDA, K-Modes) apply as used.
    Used in Eqs. 1-6 and Sec. 6.1; these are unproved background results relied on without derivation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance." pith.science (2026). https://pith.science/paper/4D6CMYOC

@misc{pith2026260808497,
  author       = {Pith},
  title        = {Pith review of: SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4D6CMYOC}},
  note         = {Machine review of arXiv:2608.08497}
}
read the original abstract

The emergence of social finance (SocialFi) transforms online communities into complex socio-economic systems. Within these spaces, collective decisions shape a "digital commons" characterized by social capital (e.g., community trust) and financial health (e.g., market liquidity). Governing such hybrid ecosystems is challenging because real-world interventions are costly and irreversible. While counterfactual simulation is essential for exploring alternative governance strategies, existing approaches fail to capture the non-linear interplay between governance rules, individual behaviors, and emergent economic outcomes. To systematically unpack this complexity, we operationalize the Institutional Analysis and Development (IAD) framework as our theoretical foundation, synthesizing prior literature with insights from formative expert interviews. Built on this framework, we present SocialFiVis, an IAD-embedded visual analytics sandbox. It introduces a robust model to quantify the dual-track digital commons, coupled with a two-phase simulation engine. This engine combines LLM-derived personas with a mechanism-guided Perception-Reasoning-Action (PRA) runtime to simulate heterogeneous, context-aware agents empirically grounded in the retained messaging cohort. A hierarchical multi-view interface with interpretable reasoning pathways enables community operators to explore counterfactual policies and trace system-level outcomes back to individual behavioral rationales. We evaluate SocialFiVis through two case studies, a user study, and follow-up interviews. Results demonstrate that SocialFiVis supports fine-grained behavioral attribution and helps explain emergent phenomena such as the structural decoupling of social capital and the resilience of messaging members under localized governance shocks.

Figures

Figures reproduced from arXiv: 2608.08497 by the authors.

Figure 1
Figure 1. IAD-embedded system overview. (1) Data Processing and Quantification (A, B) establish baseline [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The SocialFiVis interface. (A) Control Panel for case selection; (B) Persona View of persona–trait compositions, with (B1) population details on hover and (B2) a Persona Card of profile and behavioral metrics; (C) Event Timeline of (C1) macro crypto trends and milestones over community metrics, with (C2) user-injected interventions; (D) Behavior View tracking persona trajectories; and (E) Communication Network and R… view at source ↗
Figure 3
Figure 3. Behavior View workflow. (A1) Time window selection loads ground truth community dynamics (B1), while (A2) injecting a negative event [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Visual design of the Communication Network. (A) Visual encodings show persona-driven social activation via posting volumes, interaction flows, and active-agent counts. (B) Cross-view persona/date filters support focused exploration of reasoning pathways. Justification.…
Figure 5
Figure 5. Figure 5: Illustration of Case I. (A) Counterfactual simulations in the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: User study results (N=13). Q1–Q12 are rated on a 5-point [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Expert validation of the 7D persona codebook (N=10). [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 40 canonical work pages

  1. [1]

    Agresti.Categorical Data Analysis

    A. Agresti.Categorical Data Analysis. John Wiley & Sons, 3rd ed., 2013. 5

  2. [2]

    Al-Oufi, H.-N

    S. Al-Oufi, H.-N. Kim, and A. El Saddik. A group trust metric for identifying people of trust in online social networks.Expert Syst. Appl., 39(18):13173–13181, 2012. doi: 10.1016/j.eswa.2012.05.084 4

  3. [3]

    Atrey, K

    A. Atrey, K. Clary, and D. Jensen. Exploratory not explanatory: Counter- factual analysis of saliency maps for deep reinforcement learning, 2019. arXiv preprint. doi: 10.48550/arXiv.1912.05743 2

  4. [4]

    Barde and S

    S. Barde and S. van der Hoog. An empirical validation protocol for large- scale agent-based models. Technical Report 04-2017, Bielefeld University,

  5. [5]

    D. J. Berndt and J. Clifford. Using dynamic time warping to find patterns in time series. InProc. KDD Workshop, pp. 359–370. AAAI Press, Menlo Park, 1994. 5

  6. [6]

    D. M. Blei, A. Y . Ng, and M. I. Jordan. Latent Dirichlet allocation.J. Mach. Learn. Res., 3:993–1022, 2003. 4

  7. [7]

    Bonabeau

    E. Bonabeau. Agent-based modeling: Methods and techniques for simu- lating human systems.Proc. Natl. Acad. Sci., 99(Suppl. 3):7280–7287,

  8. [8]

    Borland, A

    D. Borland, A. Z. Wang, and D. Gotz. Using counterfactuals to im- prove causal inferences from visualizations.IEEE Comput. Graph. Appl., 44(1):95–104, 2024. doi: 10.1109/MCG.2023.3338788 2

Show all 89 references
  1. [9]

    Brahmstaedt

    K. Brahmstaedt. Community and consumer dynamics in NFTs: Under- standing digital asset value through social engagement.J. Consum. Behav., 24(4):1630–1655, 2025. doi: 10.1002/cb.2482 1, 2

  2. [10]

    Y . Cao, Q. Shi, L. Shen, K. Chen, Y . Wang, W. Zeng et al. NFTracer: Tracing NFT impact dynamics in transaction-flow substitutive systems with visual analytics.IEEE Trans. Vis. Comput. Graph., 31(8):4369–4386,

  3. [11]

    Y . Cao, M. Xia, K. Shigyo, F. Cheng, Q. Yu, X. Yang et al. NFTeller: Dual-centric visual analytics for assessing market performance of NFT collectibles. InProc. VINCI, art. no. 20, 8 pp. ACM, New York, 2023. doi: 10.1145/3615522.3615578 1, 7

  4. [12]

    H. Chen, C. Zhou, A. El Saddik, and W. Cai. Decentralized Web3 non- fungible token community for societal prosperity? a social capital perspec- tive.Proc. ACM Hum.-Comput. Interact., 9(2), art. no. CSCW058, 36 pp.,

  5. [13]

    L. Chen, Y . Zhang, J. Feng, H. Chai, H. Zhang, B. Fan et al. AI agent behavioral science.Humanit. Soc. Sci. Commun., 13, art. no. 1011, 2026. doi: 10.1057/s41599-026-07316-7 2

  6. [14]

    Chiu, M.-H

    C.-M. Chiu, M.-H. Hsu, and E. T. Wang. Understanding knowledge sharing in virtual communities: An integration of social capital and social cognitive theories.Decis. Support Syst., 42(3):1872–1888, 2006. doi: 10. 1016/j.dss.2006.04.001 1, 2

  7. [15]

    Y . Choi, E. J. Kang, S. Choi, M. K. Lee, and J. Kim. Proxona: Supporting creators’ sensemaking and ideation with LLM-powered audience personas. InProc. CHI, art. no. 149, 32 pp. ACM, New York, 2025. doi: 10.1145/ 3706598.3714034 1, 2, 4

  8. [16]

    doi: 10.1145/3710956 1, 2

  9. [17]

    W. G. Cochran. Theχ2 test of goodness of fit.Ann. Math. Stat., 23(3):315– 345, 1952. doi: 10.1214/aoms/1177729380 5

  10. [18]

    Elmqvist and J.-D

    N. Elmqvist and J.-D. Fekete. Hierarchical aggregation for information visualization: Overview, techniques, and design guidelines.IEEE Trans. Vis. Comput. Graph., 16(3):439–454, 2010. doi: 10.1109/TVCG.2009.84 9

  11. [19]

    D. M. Endres and J. E. Schindelin. A new metric for probability distribu- tions.IEEE Trans. Inf. Theory, 49(7):1858–1860, 2003. doi: 10.1109/TIT. 2003.813506 5

  12. [20]

    M.-L. Chu, L. Terhorst, K. Reed, T. Ni, W. Chen, and R. Lin. LLM-based multi-agent system for simulating and analyzing marketing and consumer behavior. InProc. IEEE ICEBE, pp. 72–79. IEEE, Piscataway, 2025. doi: 10.1109/ICEBE68123.2025.00018 1, 2

  13. [21]

    J. M. Epstein and R. Axtell.Growing artificial societies: social science from the bottom up. Brookings Institution Press, 1996. doi: 10.7551/ mitpress/3374.001.0001 5

  14. [22]

    Esposito, T

    M. Esposito, T. Tse, and D. Goh. Decentralizing governance: Explor- ing the dynamics and challenges of digital commons and DAOs.Front. Blockchain, 8, art. no. 1538227, 13 pp., 2025. doi: 10.3389/fbloc.2025. 1538227 1

  15. [23]

    Fagiolo, A

    G. Fagiolo, A. Moneta, and P. Windrum. A critical guide to empirical validation of agent-based models in economics: Methodologies, proce- dures, and open problems.Comput. Econ., 30(3):195–226, 2007. doi: 10. 1007/s10614-007-9104-4 2, 5, 9

  16. [24]

    J. M. Epstein. Agent-based computational models and generative social science.Complexity, 4(5):41–60, 1999. doi: 10.1002/(SICI)1099-0526 (199905/06)4:5<41::AID-CPLX9>3.0.CO;2-F 2

  17. [25]

    Gajcin and I

    J. Gajcin and I. Dusparic. Redefining counterfactual explanations for reinforcement learning: Overview, challenges and opportunities.ACM Comput. Surv., 56(9), art. no. 219, 33 pp., 2024. doi: 10.1145/3648472 2

  18. [26]

    C. Gao, X. Lan, N. Li, Y . Yuan, J. Ding, Z. Zhou et al. Large language models empowered agent-based modeling and simulation: A survey and perspectives.Humanit. Soc. Sci. Commun., 11(1), art. no. 1259, 24 pp.,

  19. [27]

    L. W. Ge, M. Easterday, M. Kay, E. Dimara, P. Cheng, and S. L. Franconeri. V-FRAMER: Visualization framework for mitigating reasoning errors in public policy. InProc. CHI, art. no. 390, 15 pp. ACM, New York, 2024. doi: 10.1145/3613904.3642750 2

  20. [28]

    Filippas, J

    A. Filippas, J. J. Horton, and B. S. Manning. Large language models as simulated economic agents: What can we learn from homo silicus? InProc. ACM EC, pp. 614–615. ACM, New York, 2024. doi: 10.1145/ 3670865.3673513 2

  21. [29]

    Guerini and A

    M. Guerini and A. Moneta. A method for agent-based models validation. J. Econ. Dyn. Control, 82:125–141, 2017. doi: 10.1016/j.jedc.2017.06. 001 2

  22. [30]

    Guidi and A

    B. Guidi and A. Michienzi. SocialFi: Towards the new shape of social media.ACM SIGWEB Newsl., 2022(Summer), art. no. 5, 8 pp., 2022. doi: 10.1145/3545196.3545201 1, 2

  23. [31]

    Ö. Gürcan. LLM-augmented agent-based modelling for social simulations: Challenges and opportunities. InProc. HHAI, pp. 134–144. IOS Press, Amsterdam, 2024. doi: 10.3233/FAIA240190 2

  24. [32]

    R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman. Metrics for explainable AI: Challenges and prospects, 2018. arXiv preprint. doi: 10. 48550/arXiv.1812.04608 8

  25. [33]

    T. L. Griffiths and M. Steyvers. Finding scientific topics.Proc. Natl. Acad. Sci., 101(Suppl. 1):5228–5235, 2004. doi: 10.1073/pnas.0307752101 4

  26. [34]

    Huang and M

    Z. Huang and M. K. Ng. A fuzzy k-modes algorithm for clustering categorical data.IEEE Trans. Fuzzy Syst., 7(4):446–452, 1999. doi: 10. 1109/91.784206 5

  27. [35]

    Hunter, E

    L. Hunter, E. Webster, and A. Wyatt. Measuring intangible capital: A review of current practice.Aust. Account. Rev., 15(36):4–21, 2005. doi: 10.1111/j.1835-2561.2005.tb00288.x 9

  28. [36]

    Imani Rad and S

    A. Imani Rad and S. Banaeian Far. SocialFi transforms social media: An overview of key technologies, challenges, and opportunities of the future generation of social media.Soc. Netw. Anal. Min., 13(1), art. no. 42, 2023. doi: 10.1007/s13278-023-01050-7 1

  29. [37]

    S. W. Jeong, S. Ha, and K.-H. Lee. How to measure social capital in an online brand community? a comparison of three social capital scales.J. Bus. Res., 131:652–663, 2021. doi: 10.1016/j.jbusres.2020.07.051 2

  30. [38]

    Z. Huang. Extensions to the k-means algorithm for clustering large data sets with categorical values.Data Min. Knowl. Discov., 2(3):283–304,

  31. [39]

    Kapoor, D

    A. Kapoor, D. Guhathakurta, M. Mathur, R. Yadav, M. Gupta, and P. Ku- maraguru. TweetBoost: Influence of social media on NFT valuation. In Proc. WWW Companion, pp. 621–629. ACM, New York, 2022. doi: 10. 1145/3487553.3524642 1

  32. [40]

    S. Kaul, D. Borland, N. Cao, and D. Gotz. Improving visualization interpretation using counterfactuals.IEEE Trans. Vis. Comput. Graph., 28(1):998–1008, 2022. doi: 10.1109/TVCG.2021.3114779 2

  33. [41]

    H. Lam, E. Bertini, P. Isenberg, C. Plaisant, and S. Carpendale. Empirical studies in information visualization: Seven scenarios.IEEE Trans. Vis. Comput. Graph., 18(9):1520–1536, 2012. doi: 10.1109/TVCG.2011.279 10 To appear in IEEE Transactions on Visualization and Computer G...

  34. [42]

    Larooij and P

    M. Larooij and P. Törnberg. Validation is the central challenge for genera- tive social simulation: A critical review of LLMs in agent-based modeling. Artif. Intell. Rev., 59(1), art. no. 15, 31 pp., 2026. doi: 10.1007/s10462 -025-11412-6 2

  35. [43]

    R. Li, S. Ye, Y . Lin, B. Zhou, Z. Kang, T.-Q. Peng et al. Causality-based visual analytics of sentiment contagion in social media topics.IEEE Trans. Vis. Comput. Graph., 32(1):35–45, 2026. doi: 10.1109/TVCG.2025 .3633839 2

  36. [44]

    Jones, S

    R. Jones, S. Sharkey, J. Smithson, T. Ford, T. Emmens, E. Hewis et al. Using metrics to describe the participative stances of members within discussion forums.J. Med. Internet Res., 13(1), art. no. e3, 2011. doi: 10. 2196/jmir.1591 4

  37. [45]

    Z. Lin, Y . Shan, L. Gao, X. Jia, and S. Chen. SimSpark: Interactive simulation of social media behaviors.Proc. ACM Hum.-Comput. Interact., 9(2), art. no. CSCW168, 32 pp., 2025. doi: 10.1145/3711066 1, 5

  38. [46]

    Macal and M

    C. Macal and M. North. Introductory tutorial: Agent-based modeling and simulation. InProc. WSC, pp. 6–20. IEEE, Piscataway, 2014. doi: 10. 1109/WSC.2014.7019874 5

  39. [47]

    C. M. Macal and M. J. North. Tutorial on agent-based modelling and simulation.J. Simul., 4(3):151–162, 2010. doi: 10.1057/jos.2010.3 2

  40. [48]

    M. D. McGinnis. Connecting commons and the IAD framework. In Routledge Handbook of the Study of the Commons, pp. 50–62. Routledge,

  41. [49]

    mfers.https://mfers.art/

    MFERS. mfers.https://mfers.art/. Accessed: 2026-03-15. 4

  42. [50]

    Li and Y

    S. Li and Y . Chen. Governing decentralized autonomous organizations as digital commons.J. Bus. Ventur. Insights, 21, art. no. e00450, 2024. doi: 10.1016/j.jbvi.2024.e00450 1

  43. [51]

    M. F. Morell. Governance of online creation communities for the building of digital commons: Viewed through the framework of institutional analy- sis and development. InGoverning Knowledge Commons, pp. 281–312. Oxford Univ. Press, 2014. doi: 10.1093/acprof:oso/9780199972036.00...

  44. [52]

    Mosqueira-Rey, E

    E. Mosqueira-Rey, E. Hernández-Pereira, D. Alonso-Ríos, J. Bobes- Bascarán, and Á. Fernández-Leal. Human-in-the-loop machine learning: A state of the art.Artif. Intell. Rev., 56(4):3005–3054, 2023. doi: 10. 1007/s10462-022-10246-w 5

  45. [53]

    Y . Ni, P. Chiang, M.-Y . Day, and Y . Chen. Using big data analytics and heatmap matrix visualization to enhance cryptocurrency trading decisions. Appl. Sci., 14(1), art. no. 154, 16 pp., 2024. doi: 10.3390/app14010154 2

  46. [54]

    J. Nielsen. The 90-9-1 rule for participation inequality in social me- dia and online communities. Nielsen Norman Group, https://www. nngroup.com/articles/participation-inequality/, 2006. Ac- cessed: 2026-06-15. 9

  47. [55]

    E. Ostrom. The institutional analysis and development framework and the commons.Cornell Law Rev., 95(4):807–816, 2010. 1, 2, 3, 9

  48. [56]

    Papakyriakopoulos, J

    O. Papakyriakopoulos, J. C. Medina Serrano, and S. Hegelich. Political communication on social media: A tale of hyperactive users and bias in recommender systems.Online Soc. Netw. Media, 15, art. no. 100058,

  49. [57]

    Mimic Shhans

    Mimic Shhans. Mimic Shhans. https://mimicshhans.com/. Ac- cessed: 2026-03-15. 4

  50. [58]

    Ricci, B

    L. Ricci, B. Guidi, A. Michienzi, A. Tagarelli, and S. Gaito. AWESOME: Analysis framework for Web3 social media. InProc. OASIS, pp. 41–47. ACM, New York, 2024. doi: 10.1145/3677117.3685010 2

  51. [59]

    Sánchez-Arrieta, R

    N. Sánchez-Arrieta, R. A. González, A. Cañabate, and F. Sabate. Social capital on social networking sites: A social network perspective.Sustain- ability, 13(9), art. no. 5147, 2021. doi: 10.3390/su13095147 1

  52. [60]

    Schlager and M

    E. Schlager and M. Cox. The IAD framework and the SES framework: An introduction and assessment of the Ostrom workshop frameworks. In Theories of the Policy Process, pp. 215–252. Routledge, 4th ed., 2018. doi: 10.4324/9780429494284-7 2

  53. [61]

    Sheng, Y

    R. Sheng, Y . Wang, X. Wang, S. Dai, Q. Guo, T.-Q. Peng et al. EMINDS: Understanding user behavior progression for mental health exploration on social media.IEEE Trans. Vis. Comput. Graph., 32(2):2284–2299, 2026. doi: 10.1109/TVCG.2025.3630646 6

  54. [62]

    Shneiderman

    B. Shneiderman. The eyes have it: A task by data type taxonomy for information visualizations. InProc. IEEE Symp. Visual Languages, pp. 336–343. IEEE, Piscataway, 1996. doi: 10.1109/VL.1996.545307 9

  55. [63]

    Sohns, C

    J.-T. Sohns, C. Garth, and H. Leitte. Decision boundary visualization for counterfactual reasoning.Comput. Graph. Forum, 42(1):7–20, 2023. doi: 10.1111/cgf.14650 2

  56. [64]

    Soroka and S

    V . Soroka and S. Rafaeli. Invisible participants: How cultural capital relates to lurking behavior. InProc. WWW, pp. 163–172. ACM, New York,

  57. [65]

    J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein. Generative agents: Interactive simulacra of human behavior. InProc. UIST, art. no. 2, 22 pp. ACM, New York, 2023. doi: 10.1145/3586183.3606763 2, 5

  58. [66]

    J. Tang, H. Gao, X. Pan, L. Wang, H. Tan, D. Gao et al. GenSim: A general social simulation platform with large language model based agents. In Proc. NAACL (System Demonstrations), pp. 143–150. ACL, Albuquerque,

  59. [67]

    Tovanich, N

    N. Tovanich, N. Heulot, J.-D. Fekete, and P. Isenberg. Visualization of blockchain data: A systematic review.IEEE Trans. Vis. Comput. Graph., 27(7):3135–3152, 2021. doi: 10.1109/TVCG.2019.2963018 2

  60. [68]

    Tovanich, N

    N. Tovanich, N. Soulié, N. Heulot, and P. Isenberg. MiningVis: Visual analytics of the Bitcoin mining economy.IEEE Trans. Vis. Comput. Graph., 28(1):868–878, 2022. doi: 10.1109/TVCG.2021.3114821 6

  61. [69]

    E. Wall, L. Matzen, M. El-Assady, P. Masters, H. Hosseinpour, A. Endert et al. Trust junk and evil knobs: Calibrating trust in AI visualization. InProc. IEEE PacificVis, pp. 22–31. IEEE, Piscataway, 2024. doi: 10. 1109/PacificVis60374.2024.00012 9

  62. [70]

    A. Z. Wang, D. Borland, and D. Gotz. An empirical study of counterfactual visualization to support visual causal inference.Inf. Vis., 23(2):197–214,

  63. [71]

    A. Z. Wang, D. Borland, and D. Gotz. A framework to improve causal inferences from visualizations using counterfactual operators.Inf. Vis., 24(1):24–41, 2025. doi: 10.1177/14738716241265120 1, 2

  64. [72]

    X. Wen, T. D. Nguyen, S. Ruan, Q. Shen, J. Sun, F. Zhu et al. PonziLens+: Visualizing bytecode actions for smart Ponzi scheme identification.IEEE Trans. Vis. Comput. Graph., 31(9):6451–6465, 2025. doi: 10.1109/TVCG. 2024.3516379 2

  65. [73]

    X. Wen, Y . Wang, X. Yue, F. Zhu, and M. Zhu. NFTDisk: Visual detection of wash trading in NFT markets. InProc. CHI, art. no. 215, 15 pp. ACM, New York, 2023. doi: 10.1145/3544548.3581466 6

  66. [74]

    Taillandier, J

    P. Taillandier, J. D. Zucker, A. Grignard, B. Gaudou, N. Q. Huynh, and A. Drogoul. Integrating LLM in agent-based social simulation: Opportu- nities and challenges, 2025. arXiv preprint. doi: 10.48550/arXiv.2507. 19364 2

  67. [75]

    Zhang, J

    X. Zhang, J. Lin, X. Mou, S. Yang, X. Liu, L. Sun et al. SocioVerse: A world model for social simulation powered by LLM agents and a pool of 10 million real-world users, 2025. arXiv preprint. doi: 10.48550/arXiv. 2504.10157 2

  68. [76]

    doi: 10.18653/v1/2025.naacl-demo.15 2

  69. [77]

    F. Zhou, Y . Chen, C. Zhu, L. Jiang, X. Liao, Z. Zhong et al. Visual analysis of money laundering in cryptocurrency exchange.IEEE Trans. Comput. Soc. Syst., 11(1):731–745, 2024. doi: 10.1109/TCSS.2022.3231687 2

  70. [78]

    Ziems, W

    C. Ziems, W. Held, O. Shaikh, J. Chen, Z. Zhang, and D. Yang. Can large language models transform computational social science?Comput. Linguist., 50(1):237–291, 2024. doi: 10.1162/coli_a_00502 2 A APPENDIX Definition Clarity Coverage Completeness Dimension Distinction Simulati...

  71. [81]

    doi: 10.1177/14738716241229437 2

  72. [85]

    Windrum, G

    P. Windrum, G. Fagiolo, and A. Moneta. Empirical validation of agent- based models: Alternatives and prospects.J. Artif. Soc. Soc. Simul., 10(2), art. no. 8, 2007. 1, 2, 5, 9

  73. [87]

    Zhong, S

    Z. Zhong, S. Wei, Y . Xu, Y . Zhao, F. Zhou, F. Luo et al. SilkViser: A visual explorer of blockchain-based cryptocurrency transaction data. In Proc. IEEE VAST, pp. 95–106. IEEE, Piscataway, 2020. doi: 10.1109/ V AST50239.2020.00014 2

  74. [1998]

    doi: 10.1023/A:1009769707641 5

  75. [2002]

    doi: 10.1073/pnas.082080899 2, 5

  76. [2006]

    doi: 10.1145/1135777.1135806 9

  77. [2017]

    doi: 10.4119/unibi/2912187 2

  78. [2019]

    doi: 10.4324/9781315162782-5 2, 3

  79. [2020]

    doi: 10.1016/j.osnem.2019.100058 4

  80. [2024]

    doi: 10.1057/s41599-024-03611-3 2, 5

  81. [2025]

    doi: 10.1109/TVCG.2024.3402834 2

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.