AI traders lose 22% on Kalshi but near break-even on Polymarket
57-day test shows platform design decides which frontier models can profit from real trades
· “Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets”
Economics
sort pith recommended most recent
57-day test shows platform design decides which frontier models can profit from real trades
· “Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets”
Balances bias from longer windows against variance from shorter ones for lower overall error in small samples.
Even with thousands of units, only a few covariates can be balanced; new designs target smaller function classes instead.
· “The Limits of Experimental Design: Covariate Balance Beyond Low Dimension”
Dynamic learning and stopping collapse to a static convex program, priced by a shadow cost of information over time.
In i.i.d. random markets, every Pareto-efficient improvement of Deferred Acceptance provably breaks the logarithmic average-rank barrier.
· “Efficiency Adjustments Break the Logarithmic Rank Barrier”
Disclosure plus delegation replaces transfers and keeps the agent's benefit private.
Seven axioms pin down a recursive utility tree that unifies joint, separate, and conditional risk evaluation.
Why care: per-family scores show routing by task type beats picking one big model.
· “EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data”
As synthetic substitutes erode middle-tier knowledge work, governance must treat provenance verification as labor infrastructure to support
· “Human-Provenance Verification should be Treated as Labor Infrastructure in AI-Saturated Markets”
Seller inference recovers budgets nearly one-for-one from natural-language profiles, and confidentiality instructions do not stop it.
· “When Agents Shop for You: Role Coherence in AI-Mediated Markets”
Field data from 713,564 prompts shows sophistication rises with rank and function, with no lasting training gains.
· “Sophistication in GenAI Use: Field Evidence from a Large Firm”
Closed-form projection replaces iterative fitting and unifies logarithmic, Saaty, and SVD windowing.
When the cross-section variance of averaged data grows with time, ordinary regressions beat root-N; mean-reverting data do not.
· “Cross-Section Estimation of Long-Run Relations Using Time-Compressed Data”
Consistent Bayesian inference for set-identified models, with no covariate binning or ad hoc moment selection.
· “Nonparametric Bayesian Inference for Partially Identified Discrete Response Models”
A tractable model of pumping, trading, and banking gives regulators a simulation lab before rules are locked in.
A pre-announced deadline triggers an adoption surge, so limited-duration rebates deliver more rooftop solar per dollar.
A revealed-preference theorem says full-horizon rankings must extend a protected partial order, giving sharp tests.
· “Directional Revision under Two-Horizon Deliberation: A Revealed-Preference Analysis”
One OLS regression per series isolates stationary cyclical components, removing the need to pre-test for unit roots or choose…
· “Principal Component Analysis for a Mix of Stationary and Nonstationary Variables”
Label a random sample, learn the model's errors, and black-box AI can serve as credible measurement for economics.
· “The Measurement Revolution? Credible Measurement and Inference in the Age of AI”
Compares reweighted distributions to test selection on observables in panels with refreshment samples.
· “Testing selection on observables in parametric models with refreshment samples”
Y02 errors are systematic: digital tech is overcounted while heavy-industry decarbonization is undercounted.
· “Systematic Bias in Green Patent Classification: Silent Green and False Green”
In contests with fixed prizes, the effort-maximizing grading policy irons out locally misordered incentive returns, pooling adjacent ranks…
Using 115 countries, the study links a nation's culture to whether its government keeps constitutional promises.
The procedure returns an error-controlled class or "inconclusive", instead of only rejecting or failing to reject a null hypothesis.
Same-size replications of p=0.05 findings usually fail: original studies were underpowered, not selectively published.
Sequential monitoring spots overconfident Bayesian forecasts during crises that static tests miss.
· “Sequentially valid inference for probabilistic inflation forecasts”
Stacked county comparison finds no comparable penalty for Democratic women; pooled nulls hid the asymmetry.
· “Female Nomination and Party Vote Share in US Gubernatorial Elections”
Tail-heaviness gaps give a clean asymmetry, and proxy-adjustment recovers it when heavy-tailed confounders interfere.
· “Identification and Inference for Causal Effects in Extremes under General Conditions”
It uses graph-weighted correlation between treatment and residual to test exposure mappings in experiments.
· “Randomization tests for model specification in causal inference under network interference”
An operator-orthogonal double machine learning estimator delivers root-n normal inference for spatial autoregressive parameters when the…
· “Double/Debiased Machine Learning for Functional-Form-Robust Spatial Autoregression”
In majority-rule team contests, finer information (more disclosure or a finer schedule) spreads each battle's incentives without changing…
· “Outcome Disclosure and Temporal Refinement in Multi-Battle Team Contests”
Decomposes drug effects into direct, transition, and path-specific pieces without sequential ignorability.
· “Estimating Pathway Treatment Effects in the Presence of Intermediate Events with Multi-State Data”
Plug-in kernel estimator skips regularization, matches the optimal nonparametric rate, and gets a uniform bootstrap band.
Projecting factor loadings onto covariates and spatial lags removes the fixed-T bias; county data highlight education.
With nondecreasing valuation densities, the best mechanism is either two posted prices or one bundle price.
A Gaussian bootstrap approximates the whole curve, no limiting degree distribution required.
· “Uniform Inference on Quantile Effects under Network Interference”
Consensus needs reciprocity between merged groups, not between pairs, so one seed dyad can absorb the whole society.
A one-step corrected estimator keeps Wald and LR tests at nominal size even on the parameter-space boundary.
· “Uniformly Valid Inference Under Interactive and High-Dimensional Constraints”
But 18 countries are locked out of debt relief and 16 out of remittance mobilisation—and they are different countries.
It rises with query fit, doubles when text is uninformative, and recruits clicks that convert less often.
· “Contextual Visual Distinctiveness in Online Product Search”
Tangent-line twisted proposals give a computable sharp bound, so exact smoothing draws stay practical at long horizons.
· “Exact Rejection Sampling for Non-Gaussian State Space Models”
A 60-day test on 8.5 million subscribers shows algorithm upgrades lift engagement and cut superstar concentration.
· “Recommendation Quality and the Concentration of Consumption: Experimental Evidence from Netflix”
Spend reacts to budgets, promotions, TV bursts, and a bidding rule; true sales effects are known weekly.
· “A Synthetic Benchmark Dataset with Endogenous Marketing Spend for Validating Marketing Mix Models”
Instead of one ROAS number, the full time series identifies adstock, saturation, and effectiveness.
· “Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments”
Standard multivariate objective functions would fail a natural unit-change invariance, so their solutions may depend on arbitrary scales.
· “Characterizations of continuous adequate objective functions for ordinal or interval scaled data”
Firms can set one number that caps how far a tracker's beliefs about a visitor can move.
· “A Privacy Budgeting Framework for Online Experimentation”
The optimum is continuous, forced to jump, or chooses to jump — a single geometric meeting decides which.
· “Monotone Allocations without Single-Crossing: When to Bunch and When to Jump”
Anonymity, neutrality, consistency, and 1-clone invariance hold together only for Plurality.
Pessimistic voters force single-valued medians; optimistic and best-worst voters permit exactly target-set rules.
A pipeline classifies joint tours, picks their mode, and assigns drivers, so HOV and fare policies can finally be tested.
· “Accounting for intra-household joint travel in agent-based transport simulations”
Automation rankings diverge from assistant rankings; leaderboards miss a separate capability.
· “CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks”
A cutoff supply-and-demand picture yields existence in large markets and approximate stability in finite samples.
· “Stable Matching with Peer-Dependent Preferences: Existence and Cutoff Characterization”
Compare treated and untreated units with the same mediator path to separate direct from indirect policy effects.
· “Difference-in-Differences Models in the Presence of Time-Varying Mediators”
With infinitely many sources, a Bayesian receiver still learns nothing from distant chains if mutation rates are only slightly uncertain.
Both depart from the standard 20-point anchor, recasting the 19/20 cluster as low-level reasoning rather than level-9.
Every absorbing state is grand at every discount factor; failure means cycling forever between non-grand states.
From personal to work use, AI prompts carry more direction; chat modes show iterative control more than APIs.
Equilibrium spending rules are computable by path-integral control, and simulated revenues track two midstream firms.
· “Bayesian Signaling and Entry Decisions under Uncertain Market Conditions”
Voting, participatory budgeting, and citizens' assemblies now demand data-grounded theory and practical algorithms.
A single reemployment-expectation question can tell pessimists who need encouragement from optimists whom true information may demotivate.
· “Biases-Informed Job Search Guidance: Characterization, Implications, and Targeting Support”
Production losses are small; the break-even access elasticity is far below existing estimates.
· “The Geography of Research: The Trade-Off Between Knowledge Production and Access”
Portfolio-style analysis of planting decisions yields the first large-scale historical risk-aversion map and links it to tractor adoption.
· “The Yeoman's Portfolio: Measuring Historical Risk Preferences Using Crop Choice”
Even mechanisms that are not strategy-proof give every agent the same best and worst houses as top trading cycles.
· “Non-obvious Manipulability with Groups in Shapley-Scarf Housing Markets”
Recursive logit, dynamic discrete choice, and maximum-entropy IRL are formally connected — yet estimate different objects.
Hajek, AIPW, and TMLE all sit on the same inverse-probability basis; the optimal basis beats them.
· “Optimal Control Variates for Survey Sampling and Causal Inference”
Why care: financial observables alone would not reveal how time feels.
· “A Neurofinance Framework for Subjective Temporal Perception, Risk, and Investment Behavior”
A theory bounds the bits a session handover needs and the regression error a memory cap costs.
· “Handover of In-Context Learning State Across Session Boundaries”
A new axiom system recovers the exact metric behind perceived complexity, so 'twice as complex' becomes meaningful.
Bayesian searchers who only see their wins keep a myopic stopping rule; full information usually destroys it.
Regime-dependent estimates for Hungary find the buffer works: release helps, build-up costs little.
· “Macroprudential Policy and Downside Risk: Regime-Dependent Effects of Capital Regulation”
When a future self may be swayed by scandal, worst-case optimal disclosure separates irrelevant states.
Economists at universities most reliant on US federal money cut politically flagged wording after 2025.
Co-owned outlet pairs also report tone 0.342 more alike, even after fixed effects.
A 500-journal census shows fees are a residual price on top of waiting time, not a cure for congestion.
A sender who fears extra information is modeled by a leakage set; new axioms make that model testable.
Long-term, dilutable interbank debt implements the optimal triage of shocked banks with no regulator and no bailout.
· “Systemic Risk in Financial Networks Revisited: Debt Dilution as a Backdoor Bail-in”
Identified shocks serve as instruments, so misspecification elsewhere cannot contaminate block-level estimates
· “Limited-Information Estimation of Heterogeneous Agent Models”
Gains went to the bottom half of the wage distribution, opposite the usual skill-biased pattern.
· “Financial Technologies, Labor Markets, and Wage Inequality: Evidence from Instant Payment Systems”
The paper proves the averaged stochastic-score estimator inherits the MLE's limiting distribution and sandwich inference.
· “Scalable likelihood-based inference for limited dependent variable models”
A Bernoulli design with nonlinear shrinkage is the optimal procedure, beating complete randomization for small trials.