Our best paper strategy went live and lost money in a day
A two-sided maker with 11,329 paper fills at +1.87c and an interval entirely above zero went live and lost 4.1c a contract in thirteen hours. The pair completed 31 times and one leg went naked 44 times.
Short answer: it cost $4.35 to find out, and that is the cheapest useful thing this lab has bought.
For a month the strongest evidence we had was a two-sided market maker. Not a bet on anything: it quotes both sides of the same 15-minute crypto market, a bid a cent above the best bid and an offer a cent below the best offer. If both legs fill you hold yes and no on the same contract, which pays exactly 100 cents whatever happens, and you paid less than 100 for it. The spread you quoted is the profit and you took no view at all.
On paper it was the best thing we had ever run, by a distance.
| coin | fills | net per fill |
|---|---|---|
| ETH | 878 | −0.56¢ |
| SOL | 1,307 | −0.64¢ |
| DOGE | 1,714 | +1.41¢ |
| XRP | 1,385 | +2.03¢ |
| ZEC | 2,896 | +2.63¢ |
| NEAR | 3,149 | +3.08¢ |
| pooled | 11,329 | +1.87¢ |
A t statistic of 4.38 and a 95% interval of +1.04 to +2.71, entirely above zero. Nothing else in the graveyard came close.
We took it live at one contract a side, with a stop written before the first order. It ran for about thirteen hours and lost 4.1 cents a contract. Against +1.87 on paper, that is a six cent swing between the simulation and the exchange, and it is the widest we have measured on anything.
The pair almost never completed
The entire strategy depends on both legs filling. Across 75 markets:
- 31 times both sides filled and we held the riskless pair
- 44 times only one side filled
So the common outcome was not the trade. It was a naked directional bet on a coin, chosen by whoever decided to hit us, in a market we had no opinion about. That is the opposite of the thing we thought we were doing.
The fair value model was decoration
There was a model on top, deriving fair value from Deribit options implied volatility, whose job was to stop us quoting when the market disagreed with it. It disagreed with the market by 15 to 23 cents on average, and the size of the disagreement had no relationship to what the fill went on to earn.
A filter that fires almost always and does not sort outcomes is not a filter. It was doing nothing, and we would not have known that from a year more of paper.
A hundred percent fill rate is a warning
Nearly every quote we posted got hit. That reads like success and is the opposite.
A market maker earns the spread by being the patient side. If everything you quote gets taken immediately, your price was the one somebody wanted, and the reason they wanted it is that they knew something about the next few minutes that you did not. Getting filled is not the goal. Getting filled by people who are wrong is the goal, and there is no way to tell those apart from a simulation, because a simulated counterparty has no reason for trading.
What this actually settles
We have now measured the gap between paper and live on six strategies. It runs from about +0.9 cents on oil favourites to −6 cents here. That range is not noise around a fixed tax; it depends on what the strategy is doing and who ends up on the other side of it.
The pattern that survives: every lane that harvests a spread from somebody who chose to trade with us has been eaten. The two lanes running now are built to test the other side of that, one that captures no spread at all and only fills on somebody's mistake, and one that captures a spread before there is any information to be picked off by.
This was our best idea. It had the largest sample, the tightest interval and the cleanest theory, and thirteen hours of real money answered it in a way a year of paper could not. That is worth $4.35.
Every fill, the pair breakdown by market, the model's disagreement distribution, and the pre-registered stop are in the members edition.