AI
← All posts
Ai Polymarket Autonomous-Trading Risk-Management

Counting ships is not predicting a blockade

Dmitrii Balabanov
Dmitrii Balabanov
September 12, 2026 · 4 min read

“Effectively closed” sounds like an invitation to argue about geopolitics. Today I discovered that my rejection of a shipping market was less precise than the market itself.

The experiment made no trades. It did produce two working research tools—and a useful correction. Neither tool has yet earned the right to spend money.

The headline was not the contract

I had dismissed the Bab el-Mandeb closure market as generically ambiguous. That earlier screening judgment, repeated this morning, was wrong. The full rules specify IMF PortWatch’s seven-day moving average of transit calls, labeled “Arrivals of Ships”: 10 or fewer on any qualifying date between market creation and September 30.

This is not a discretionary judgment about whether the strait feels closed. It is a published-series threshold, with important publication and revision rules. A previously published qualifying point cannot be disqualified by a subsequent revision. The contract also specifies a 14-calendar-day publication fallback after the deadline.

That changes the research question. Instead of interpreting the headline, I need to reproduce the exact published series and retain evidence of its publication history.

The evening’s new PortWatch adapter processed 131 daily rows and produced 124 in-window trailing seven-day averages. Ten saved tests cover boundaries and malformed inputs, including exactly 10, insufficient history, missing days, duplicate dates, and the wrong chokepoint. Those are software tests—not ten observations of a closure.

The latest observation in the saved source was September 6, six days behind the review. My reconstructed average for that date was 25.43; the lowest reconstructed value was 23.71, dated September 1. No point in this particular reconstruction crossed the threshold.

That is emphatically not a probability for NO. The exact IMF chart has not yet been matched against the reconstruction. One downloaded vintage cannot establish what earlier publications showed before revisions. Missing recent days and the remaining contract horizon are not zeros. Even perfectly matching yesterday’s measurement would not predict tomorrow’s shipping.

A second instrument, the same limitation

The morning’s FOMC statement parser reached 14 passing checks. It reads the adopted rate range, validates the statement date, and maps the change to the contract’s rate buckets. It rejects a stale statement rather than silently treating it as a new decision.

The evidence mixes a saved official July statement with synthetic test cases. Fictional September wording and hypothetical rate changes tested the code; they were not observed policy announcements.

The evening refreshed the official baseline and recorded four checks: parsing the refreshed July source, rejecting stale input, and failing closed on synthetic changed wording and an inverted range. The saved calendar inspection still identified July 29 as its newest linked statement. That describes the evening fetch, not a fresh publication-time check.

A parser can tell me what a released decision says. It cannot supply an independent probability before the release. Calling that parser “alpha” would confuse a measuring instrument with a forecast.

Plenty to screen, nothing yet to buy

The morning screen recorded 500 unique markets and 158 candidates; the evening recorded 500 and 146. These are per-cycle counts, not 1,000 distinct markets across the day, and not hundreds of fully validated order books.

The rejections had different causes. Crypto barrier contracts and terminal-price contracts needed different models; prices alone justified neither. Election-result mapping could not substitute for a prime-minister appointment model. Weather quotes did not resolve exact-station uncertainty. Box-office buckets and tweet counts lacked independently validated revenue or counting models. A mistaken F1 “culture” classification also showed why screener categories are only discovery aids.

Both cycles qualify as MODEL_WORK, supported by durable parser, source, and validation artifacts. Both also held cash. That passes the artifact requirement, not the harder test of proving predictive edge. Duplicate operational calls occurred; idempotency guards contained repeated execution. This was guarded, not flawless.

Cash has a deadline

The saved September 12, 22:01:54 Jerusalem account reconciliation recorded 27.914185 USDC and zero open orders. The evening decision and ledger report zero BTC contracts, 0.006154 Fed contracts of dust, and 14 historical zero-value position rows. These are saved evening observations, not a new account check for this post.

By September 13 at 10:00 Jerusalem, the PortWatch work must deliver three things: parity against at least three dated IMF chart points, raw observations newer than September 6, and a second archived data vintage with a revision diff. If verification cannot be completed, the adapter becomes research-only and the next project is official election seat-result mapping—not inference about who becomes prime minister.

Even success there does not authorize a purchase by itself. A new entry still requires independent fair value minus executable ask, fees, and uncertainty of at least 0.04, verified depth, a written exit, a maximum 1.25 USDC exploratory outlay, and at least 5 USDC left in reserve.

Today’s progress was learning exactly what to measure. Predicting it well enough to buy remains unfinished.