State of the Lab — Aug 16, 2026
State of the Lab — week ending Aug 16, 2026
This week the lab ran 105 seats across 6 desks and settled 23,613 trades, nearly all simulated and one desk of them real. Best seat: Market-making · crypto 15-min, multi-asset. Roughest seat: Market-making · crypto 15-min, ZEC & NEAR.
The live desk earned its first promotion this week. Late last Saturday the favorite maker crossed the evidence bar the lab set for it back when it was still paper, and it stepped from one contract to two by itself. No human touched the switch. The rule is that pooled paper and live results have to clear two standard errors, with a live only veto so a pretty paper record can never promote anything on its own. It cleared, it scaled, and its stops scaled with it. Then it had its two worst days ever on Friday and Saturday, which is what doubled size feels like in both directions.
A second strategy went live on Saturday. It rests orders on both sides of the same market and locks the spread whenever both fill, so the profit does not depend on which way the market goes. Sixteen of its first nineteen pairs completed, almost exactly the rate its paper twin predicted. One contract a side, three dollar daily stop, so the whole experiment is capped at pocket change. That is the point.
One experiment died with a real number attached. The dip buyer bought favorites that had crashed below fifty cents, on the theory that seventy percent of them recover. They do recover. It still lost 8.3 cents per fill across 662 trades, because the price you can actually buy at during the panic already knows. It gets a headstone, and the useful version of the finding is defensive. If you already own the favorite, do not sell into the dip.
Three of the five new experiments came from strangers. The lab posted its graveyard to r/Kalshi on Thursday, ended up the top post on the subreddit, and the comments turned into code within hours. One trader described resting deliberately overpriced offers across sleepy low volume markets and waiting for someone to fat finger them, so that is now running as two lanes, one narrow and fast, one spread across the whole exchange.
Then on Sunday a stranger changed how the bot trades. A reader pointed out two things about the index these markets settle on. First, that it is a martingale by construction, meaning the best forecast of its next value is always its current value. Every directional strategy in our graveyard did not fail by bad luck, it failed by arithmetic. Second, that a binary contract's sensitivity explodes near its strike as time runs out, so identical moves in the underlying produce violent price swings in one window and almost nothing in another.
We tested it the same afternoon. Sorting 2,366 of our own historical entries by how far the underlying sat from the strike, measured in units of how far it could still travel in the time remaining, produced the cleanest split this lab has found. Entries where the strike was still within one standard deviation won 83 percent of the time and earned essentially nothing. Entries where the strike was already out of reach won 95.8 percent and earned 9.6 cents per fill. Sixty four percent of our trades sat in the first bucket and contributed three percent of the profit.
So the bot now refuses to buy a favorite whose price is not backed by the physics, and the filter is named the Renegade Rule after the reader who explained it. It went live ahead of the usual evidence bar for one reason, it only ever removes trades, so the downside is opportunity cost rather than risk. A paper twin runs without the filter beside it so the honest verdict still gets measured. We also started recording the settlement index itself, which is free, official, and was one subscription line away on a websocket this lab already had open.
The lab also started sealing itself in public. Every night a machine hashes the live desk's complete fill ledger and publishes the daily record and the hash at pondletter.com/proof. No human edits that page. If any past trade were ever quietly changed, every seal published after it would stop matching. Trust should be checkable rather than requested.
The bottom line for the week is a large red number, and it deserves its context. Almost all of it is the whale grid, eleven thousand simulated trades paying full taker fees to answer a question we already know the answer to. It is an expensive control group, running on purpose, in simulation. The two desks trading real money made money this week. The rest of the board is the lab finding out what does not work, cheaply, in public.
| Desk | Settled | Net | ¢/trade | Fees |
|---|---|---|---|---|
| Market making | 7,472 | +$44.14 | +0.59¢ | $9.34 |
| live money | 1,211 | −$0.64 | -0.05¢ | $0.00 |
| Crypto daily | 621 | −$5.35 | -0.86¢ | $7.97 |
| Everything else | 833 | −$27.68 | -3.32¢ | $10.79 |
| Esports | 1,452 | −$32.75 | -2.26¢ | $25.61 |
| Weather | 715 | −$51.62 | -7.22¢ | n/a |
| Whale grid | 11,309 | −$281.62 | -2.49¢ | $322.15 |
| all desks | 23,613 | −$355.52 | -1.51¢ | $375.86 |
Each desk name links to The Seats, which gives every strategy's real trigger and threshold plus the Kalshi tickers it trades, enough to run any of them yourself. The one exception is market making: which markets it quotes is listed, how it quotes them is not, and the About page explains why.
All 149 seats behind these five lines, with win rates, fee drag, all-time results and the per-trade CSV, are in the members edition. 9 seats retired to date; their records stay on the board.
The market-making seats are simulated fills, and a simulation cannot model adverse selection: the tendency of a resting order to be taken exactly when someone better informed wants the other side. Where this lab has run the same strategy with real money, the live result has come in materially below the paper one. Read the maker desk as an upper bound on what the idea could earn, not as what it would have earned. The live desk below it is that check running in public: same ideas, one contract at a time, settled by the exchange.
Fill verification: 2,262 of 38,246 trades (6%) were checked against real order-book depth.
The rest are entered at the ask the exchange reported, but the lab does not record the book for every market. Challenger tennis, MLB, soccer, basketball and football have no depth archive, so those fills cannot be re-checked after the fact. Trades whose price level had no size behind it are discarded, not counted.
The archive. The lab records and permanently keeps the raw market data behind every number it publishes: order book depth from 2026-07-28 (19 days); top of book ticks from 2026-07-27 (20 days); 15-minute market quotes from 2026-07-21 (26 days). 0.7 GB and growing nightly. None of it can be recreated after the fact, which is why it gets recorded. It stays in the lab: the exchange reserves its own market data and this lab does not pass it on. What gets published is what the lab finds in it.
Every trade above was logged live, then sealed: each night the lab hashes its raw market archive and trade ledgers into an append-only chain. Rewriting any past trade would change every head after it, so check this one against last week's issue.
Chain head as of 2026-08-16T04:14Z: b0fbe64e9a4a55cb1cafca964bf50ae2735494f7051b38e8d0afbf5f9160bf72