Pith. sign in

receipts

The numbers, with their receipts

Two kinds of figure appear on this page and they are labeled. A live number comes from a database query that runs when the page renders (cached one hour) and says "as of". A measured number comes from a frozen experiment and carries the date it was run. Nothing here is a projection.

Operating totals

figurevaluereceipt
reviews published 191,649 live · count of current verdicts · as of 2026-08-05
papers cataloged 2,907,120 live · count of the papers table · as of 2026-08-05

Reviews per day, last 14 days

dayreviewstotal costavg cost per review
2026-08-05 5,778 $150.78 $0.026
2026-08-04 8,606 $243.01 $0.028
2026-08-03 9,082 $294.27 $0.032
2026-08-02 8,315 $267.11 $0.032
2026-08-01 9,446 $297.97 $0.032
2026-07-31 2,085 $607.02 $0.291
2026-07-30 1,149 $398.69 $0.347
2026-07-29 1 $0.08 $0.085

Live · verdict rows grouped by creation day, cost summed from per-review API billing. Days that mix review models mix their costs; the model in service since 2026-07-31 is DeepSeek V4 Flash at roughly three cents a review.

Verdict distribution, whole record

verdictcount
UNVERDICTED127,176
CONDITIONAL50,824
ACCEPT9,471
REJECT4,178

Live · every paper's current verdict, counted · as of 2026-08-05.

Planted-defect catch rates, measured 2026-07-31

modelplanted defects caughtnote
GPT-5.6 Luna6 of 6highest catch rate, higher cost
DeepSeek V4 Flash, with screening clause5 of 6the production configuration
DeepSeek V4 Flash, base3 of 6noticed defects but excused them; fixed by one prompt clause
Grok 4.5 high2 of 6previous production model

Method: six papers seeded with known defects, reviewed by each model under the same frozen rubric, then judged blind. The judge saw anonymized review packets and a sealed model-to-packet mapping, unsealed only after the final judgment. The screening-clause fix was validated the same way against the same frozen corpus, at zero measured noise cost. Measured 2026-07-31.

Reference-extraction yield by arXiv era, measured 2026-08-04

submission eraextraction success
1991-199996%
2000-200997%
2010-201999%
2020-202699%

Method: 800 papers drawn at random, 200 per era, run through the deterministic reference-extraction path only (no model calls). Success means at least one structured reference row extracted. This is the gate that says the historical backfill will carry its citation graph rather than bare verdicts. Measured 2026-08-04.

More: what Pith is · the scoring rubric · integrity findings · the API.