ComboCrusher

Track record

The sample is too small to mean anything.

This page will eventually be the only part of the site that matters. Right now it is mostly empty, and publishing it empty is deliberate — the shape of the record is a commitment made before there are results to be selective about.

CurrentAs of [date]

Where the sample stands.

MeasureValue
Pre-recorded estimates logged[n]
Resolved and graded[n]
Sample needed for a calibration claim~400 resolved
Brier score vs. market baselineinsufficient sample
Net ROI after feesinsufficient sample
Realized vs. modeled edgeinsufficient sample

Every estimate is written to an append-only log with a timestamp before the event resolves. Nothing is added to the record retroactively, and nothing is removed from it.

MetricsWhat gets reported

Three numbers, none of them win rate.

Calibration

Of everything estimated at 30%, roughly 30% should happen. This is plotted against the diagonal with confidence bands, and it is the single most informative thing about whether a probability model is any good. It is also the number that is hardest to fake, because being wrong in a consistent direction shows up immediately.

Net ROI after fees

Gross returns on Kalshi are a fiction. The fee schedule is nonlinear in price and scales with contract count, so a position that looks profitable at mid can be negative once it is actually filled and settled. Reported returns are net, or they are not reported.

Realized versus modeled edge

The scanner claims a specific gap in cents. The relevant question is not whether the position won, but whether the gap that was actually captured matched the gap that was predicted. A model can be profitable and still be badly wrong about why.

Why win rate is excluded

On binary contracts priced anywhere from 1¢ to 99¢, win rate is a statement about which prices were selected, not about skill. A 90% win rate on contracts bought at 95¢ is a losing operation. Anyone leading with win rate on a market like this is either confused or counting on you to be.

GradingThe rules, fixed in advance

How entries get graded.

  1. An estimate is logged with a timestamp, the market, the quoted price, the modeled fair value, and the reasoning — all before resolution.
  2. The quoted price recorded is the price observable at the moment of scoring, with the bid-ask spread recorded alongside it. Mid is used for measurement, never as an assumed fill.
  3. Entries flagged stale at alert time stay in the log and are graded like everything else. Excluding them would be exactly the kind of selection that makes records worthless.
  4. Grading is by contract resolution as published by the exchange. No discretionary voids, no "would have won if."
  5. Errors are corrected by appending a correction, never by editing history. Corrections are visible in the log.

If a rule here needs to change, the change is dated and the old rule stays visible, with the sample split at the boundary.

Build logLatest entries

What broke, and what it changed.

Every live session produces a dated entry: what was observed, and what it changed about the product. The full log goes out to the list. A sample of what these look like:

DateObservationProduct impact
[date] Same leg quoted at different prices depending on which legs it was combined with. Marginal pricing cannot be assumed consistent within a ticket. Scoring now compares against the quoted combo, never a reconstructed one.
[date] Quote probe design was structurally similar to placing orders never meant to fill. Rebuilt as read-only public order-book requests. No order creation anywhere in the pipeline.
[date] A ticker oscillated 20–35 points inside an hour with no corresponding news. Escalated to the exchange with the full trade tape. Anomaly detection added upstream of scoring.
[date] Moneyline and total for the same game were never scored together. Cross-series game key added in transform. The premise did not work at all before this.

Replace these with real dated entries from the product log before launch — the specifics are the credibility.

StandingWhat this is worth today

Assume no edge until the record says otherwise.

A few dozen results is noise. A hundred is a suggestion. Several hundred resolved positions, calibrated and net of fees, is the first point at which anyone — including whoever built this — should update much.

Until that threshold is reached, the correct interpretation of everything on this site is: here is an interesting measurement apparatus and an unfinished experiment. Not: here is an edge.

Get the log as it fills in →