Paper trading vs real money: the same bot, run both ways
One market making bot, run simultaneously on paper and with real dollars on Kalshi. The gap between the two columns is the true cost of simulation, printed weekly.
Every paper trading result you have ever seen comes with the same
asterisk: simulated fills are too kind. A simulated resting order is
always filled on its own terms. A real one is filled precisely when
someone on the other side wants it filled, which is often exactly when
you wish it had not been. The polite name for this is adverse
selection, and no simulation can price it, because the information that
causes it never touches the simulator.
So this lab ran the experiment directly. One market making book, the
best behaved of the simulated desk, runs on paper. An identical twin
runs on the exchange with real dollars, one contract at a time. Same
rules, same markets, same weeks. They appear on the weekly board as
book A and its live twin. The only difference between them is that one
is imaginary. Whatever gap opens between their results is the true
price of simulation, measured rather than warned about.
The score so far
| Paper control | Live twin | |
|---|---|---|
| Settled | 1,408 trades | 1,091 fills |
| Win rate | 83% | 84% |
| Net | $20.99 | $33.10 |
| Net per settle | +1.49¢ | +3.03¢ |
The live book has been running since 2026-08-04, and the fills are
the exchange's own records, not the lab's opinion of what happened.
The dollar figures are small because the size is deliberately small:
one contract per quote while the experiment earns trust.
The honest headline is that so far the live twin is holding its own,
which is not what this lab's own caution predicted. Earlier live runs
of other strategies came in well below their paper versions, that
warning sits under the board every single week, and a few hundred
fills of agreement do not retire it. Adverse selection does not show
up on a schedule. It shows up on the worst day, all at once. That is
precisely why this page exists: if the gap opens, it prints here.
What real money said before this
This is not the lab's first real dollar. It is roughly its two
thousandth, and the earlier ones mostly said no.
Before the desk quoted anything, the lab ran 1,656 real trades
across a series of taker bots, strategies that cross the spread and
pay the fee. Together they lost $55.19. Not catastrophically,
just steadily, the fee grinding away every marginal read, which is the
finding the fee study later put numbers on. The
survivors of that era were retired, and the lesson bought with that
money reshaped the whole lab: stop paying the toll, start collecting
it.
Four other maker books also got short live probes, 900 fills in
total, which together cost $5.41. Small samples, quickly ended,
kept on the record because the record is the product.
Add up every real fill this lab has ever made and the total stands at
-$27.50. That figure spans every live experiment from the
beginning, the retired taker bots above included, which is why it sits
below the running total on the proof page: that page seals a
single lane from 4 August onward and counts nothing before it. Both
numbers are correct and they are counting different things. This one
updates as this page does, and there is no version of this publication
where it gets edited out if it turns ugly.
What this does not prove
A few hundred live fills at one contract of size prove nothing about
capacity. Quoting one lot and quoting fifty are different jobs, and
the market treats them differently. The live twin also runs in the
lab's friendliest venue, the crypto 15 minute markets, where makers
pay no fees. And the paper control it is measured against is itself
only weeks old. This is an experiment in public, not a verdict. The
board prints both columns every Sunday, and the gap between them,
whichever way it moves, is the finding.
What each book quotes is public and how it quotes is not; the
reasoning is on the About page. The weekly board, with the
live desk marked, is in
State of the Lab.