Generative agents with memory streams, reflection, and planning using LLMs exhibit believable individual and emergent social behaviors in a simulated town.
Journal of the American Statistical Association47(260), 583–621 (1952)
16 Pith papers cite this work, alongside 12,017 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
TS-Neyman uses posterior sampling of stratum variances to implement an adaptive Neyman allocation rule that converges almost surely to the oracle proportions and achieves near-oracle efficiency in finite-strata settings.
EVENT5Ws is a new large-scale, manually verified open-domain event extraction dataset that benchmarks LLMs and demonstrates cross-context generalization.
SCARFACE provides a harmonized annual panel dataset of over 2,700 socio-economic, environmental, and agricultural indicators for 256 sub-regions in the Po Valley, Italy, spanning 2011 to 2024.
Scoring the same cervical-spine segmentations against silver rather than expert labels overestimates Dice by ~8 points and turns a non-significant age fairness gap into a significant one via variance collapse.
Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.
Derives a closed-form task-specific strictly proper scoring rule for ATE estimation by matching local curvature of the IPW error metric.
OPAL learns optimal smooth labeling policies from ML uncertainty scores to enable low-variance prediction-assisted inference with finite-sample coverage guarantees.
ExPerT infers query-specific user expertise from semantic text and keystroke dynamics via LLM prompting to adapt response generation, cutting inference error 65.7% and raising satisfaction 17.52% in a 40-participant study.
Quantum kernel methods show no statistically significant edge over strong classical baselines on tabular classification tasks, with current feature maps failing to match the spectral properties of the best classical kernel.
Case study of 18,020 Kubernetes PRs shows label-diff congruence is prevalent and stable, with higher congruence linked to fewer review participants among core developers and more among one-time contributors.
Persistent homology signatures of the solid phase in synthetic porous media correlate with diffusion-release regimes even after stratifying by target porosity, and classify early-fast, late-release, and long-tail behavior with 0.64–0.76 test accuracy.
Develops unified formulas for conditional quantities and transportation functionals via distributional derivatives and copulas, yielding quantile representations for Wasserstein distance and applications to normal approximation of counting distributions.
MODEE is a multimodal system that integrates graphs with LLM embeddings to outperform prior open-domain event extraction methods on large datasets.
Self-reported LLM usage frequency associates more consistently with pre-instruction AI perceptions than prior education or self-rated familiarity in graduate trainees.
citing papers explorer
-
Generative Agents: Interactive Simulacra of Human Behavior
Generative agents with memory streams, reflection, and planning using LLMs exhibit believable individual and emergent social behaviors in a simulated town.
-
TS-Neyman: Posterior Sampling for Adaptive Stratified Estimation
TS-Neyman uses posterior sampling of stratum variances to implement an adaptive Neyman allocation rule that converges almost surely to the oracle proportions and achieves near-oracle efficiency in finite-strata settings.
-
EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents
EVENT5Ws is a new large-scale, manually verified open-domain event extraction dataset that benchmarks LLMs and demonstrates cross-context generalization.
-
SCARFACE: a harmonized spatio-temporal dataset integrating socio-economic, environmental, and agricultural indicators for the Po Valley (Italy), 2011--2024
SCARFACE provides a harmonized annual panel dataset of over 2,700 socio-economic, environmental, and agricultural indicators for 256 sub-regions in the Po Valley, Italy, spanning 2011 to 2024.
-
False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation
Scoring the same cervical-spine segmentations against silver rather than expert labels overestimates Dice by ~8 points and turns a non-significant age fairness gap into a significant one via variance collapse.
-
Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.
-
Tailoring Strictly Proper Scoring Rules for Downstream Tasks: An Application to Causal Inference
Derives a closed-form task-specific strictly proper scoring rule for ATE estimation by matching local curvature of the IPW error metric.
-
Optimized Labeling Resource Allocation for Prediction-Assisted Inference via OPAL
OPAL learns optimal smooth labeling policies from ML uncertainty scores to enable low-variance prediction-assisted inference with finite-sample coverage guarantees.
-
ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues
ExPerT infers query-specific user expertise from semantic text and keystroke dynamics via LLM prompting to adapt response generation, cutting inference error 65.7% and raising satisfaction 17.52% in a 40-participant study.
-
Benchmarking Quantum Kernel Support Vector Machines Against Classical Baselines on Tabular Data: A Rigorous Empirical Study with Hardware Validation
Quantum kernel methods show no statistically significant edge over strong classical baselines on tabular classification tasks, with current feature maps failing to match the spectral properties of the best classical kernel.
-
Efficiency for Experts, Visibility for Newcomers: A Case Study of Label-Code Alignment in Kubernetes
Case study of 18,020 Kubernetes PRs shows label-diff congruence is prevalent and stable, with higher congruence linked to fewer review participants among core developers and more among one-time contributors.
-
Topological Signatures of Diffusive Release in Porous Media
Persistent homology signatures of the solid phase in synthetic porous media correlate with diffusion-release regimes even after stratifying by target porosity, and classify early-fast, late-release, and long-tail behavior with 0.64–0.76 test accuracy.
-
Unified formulas for conditional quantities and transportation functionals
Develops unified formulas for conditional quantities and transportation functionals via distributional derivatives and copulas, yielding quantile representations for Wasserstein distance and applications to normal approximation of counting distributions.
-
A Multimodal Text- and Graph-Based Approach for Open-Domain Event Extraction from Documents
MODEE is a multimodal system that integrates graphs with LLM embeddings to outperform prior open-domain event extraction methods on large datasets.
-
Engagement Intensity as a Learner-Modeling Signal for Adaptive AI Ethics Instruction
Self-reported LLM usage frequency associates more consistently with pre-instruction AI perceptions than prior education or self-rated familiarity in graduate trainees.
- A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation Methods