Two kinds of figure appear on this page and they are labeled. A live number
comes from a database query that runs when the page renders (cached one hour) and says
"as of". A measured number comes from a frozen experiment and carries the date it was run.
Nothing here is a projection.
Operating totals
figure
value
receipt
reviews published
191,649
live · count of current verdicts · as of 2026-08-05
papers cataloged
2,907,120
live · count of the papers table · as of 2026-08-05
Reviews per day, last 14 days
day
reviews
total cost
avg cost per review
2026-08-05
5,778
$150.78
$0.026
2026-08-04
8,606
$243.01
$0.028
2026-08-03
9,082
$294.27
$0.032
2026-08-02
8,315
$267.11
$0.032
2026-08-01
9,446
$297.97
$0.032
2026-07-31
2,085
$607.02
$0.291
2026-07-30
1,149
$398.69
$0.347
2026-07-29
1
$0.08
$0.085
Live · verdict rows grouped by creation day, cost summed from per-review API
billing. Days that mix review models mix their costs; the model in service since 2026-07-31
is DeepSeek V4 Flash at roughly three cents a review.
Verdict distribution, whole record
verdict
count
UNVERDICTED
127,176
CONDITIONAL
50,824
ACCEPT
9,471
REJECT
4,178
Live · every paper's current verdict, counted · as of 2026-08-05.
Planted-defect catch rates, measured 2026-07-31
model
planted defects caught
note
GPT-5.6 Luna
6 of 6
highest catch rate, higher cost
DeepSeek V4 Flash, with screening clause
5 of 6
the production configuration
DeepSeek V4 Flash, base
3 of 6
noticed defects but excused them; fixed by one prompt clause
Grok 4.5 high
2 of 6
previous production model
Method: six papers seeded with known defects, reviewed by each model under the
same frozen rubric, then judged blind. The judge saw anonymized review packets and a sealed
model-to-packet mapping, unsealed only after the final judgment. The screening-clause fix was
validated the same way against the same frozen corpus, at zero measured noise cost. Measured
2026-07-31.
Reference-extraction yield by arXiv era, measured 2026-08-04
submission era
extraction success
1991-1999
96%
2000-2009
97%
2010-2019
99%
2020-2026
99%
Method: 800 papers drawn at random, 200 per era, run through the deterministic
reference-extraction path only (no model calls). Success means at least one structured
reference row extracted. This is the gate that says the historical backfill will carry its
citation graph rather than bare verdicts. Measured 2026-08-04.