In which sixty-seven investigations mostly conclude "change nothing", and I argue that is the product.
The research folder currently holds 67 studies. The decision tracker holds 86 items. If you sampled their conclusions at random you would find, overwhelmingly, the same verdict:
No change to the live book.
This is not failure. It is the entire output. Let me defend that, because it is the least intuitive thing in this series and the most transferable.
Given a capable agent, the cost of asking a question of your data collapses. "What if the wings were wider?" used to be a day's work. It is now twenty minutes, and - crucially - the agent is never bored by the forty-third variation, never rounds a number to make a story tidier, and will write its own regression tests without being nagged.
So you ask everything. Wing width, swept seven ways. Position sizing per day-to-expiry. Eight different breach triggers. An entirely different structure. A direction filter with nine lookback windows. Each returns a tidy table.
And that is precisely where it becomes dangerous, because a variant that beats the baseline is the normal outcome of searching hard enough. With enough knobs, something always wins. The skill is no longer generating results. It is refusing most of them.
The regime-transfer test. The historical data spans two eras - the exchange moved the weekly expiry day partway through. Any change is measured on both. A recent candidate improved total profit across all eras and lost money in the live era. Eight previous ideas had the same signature. Verdict: no change. If it only works in conditions you no longer trade in, it is archaeology.
The tail-dependence test. For any variant, what share of total profit comes from its best 5% of trades? The live book: 33%. Several attractive-looking variants: over 100% - meaning the other 95% of their trades collectively lose money and the entire result rests on a handful of lucky weeks. One such variant turned negative when its best five weeks were removed.
The leverage test. A change that adds profit by taking more risk is not an improvement, it is a decision to bet bigger, and should be argued on those terms. One sizing variant produced 17% more profit - and its marginal return-to-drawdown ratio was 2.38 against the book's average of 22.4. Every extra unit was nine times less efficient than the book it joined. Rejected.
Applying these consistently killed eleven separate exit-rule improvements, and produced the finding that justified all of them: every rule that shortens a winning position's life loses money in proportion to the time it removes. The book's return is the premium it collects by carrying risk to expiry. Any rule that stops carrying stops earning.
I would not have found that by succeeding. I found it by failing eleven times in a consistent direction.
Honesty demands the other column. During this phase the AI - with me watching - produced:
response. Corrected, then wrong again by 3× in the opposite direction, from reading a second wrong field. The right answer required an actual observed account reading.
should have been positive - reported confidently, with an explanation attached, and only unravelled when I noticed the number moved the wrong way as a parameter changed.
had a detail section. It simply was not in the total.
None of these were caught by tests. All were caught by a human looking at a number and saying that cannot be right. The most useful question I asked all month was four words long: "does that total correctly?"
Measure your research programme by conclusions you can trust, not experiments you can run. The second is now nearly free; the first is not, and the gap between them is where credibility is won or lost.
Concretely: before you accept a result, ask whether it survives a regime it was not fitted to, whether it depends on a handful of observations, and whether it is edge or simply leverage. Three questions. They cost minutes and they will reject most of what crosses your desk - which is correct, because most of what crosses your desk is noise wearing a p-value.
This series was written with Claude, the same tool it describes. I supplied the project, the judgement calls and the arguments; Claude supplied the drafting, and dug every figure out of the repository’s own commit history so I could not flatter myself from memory.
Declaring that seems the least I can do given the subject. It would be a peculiar hypocrisy to publish eight posts on harnessing AI while implying I typed them all by hand. If the writing is good, that is partly the tool. If the judgement is sound, that part is mine. Distinguishing between those two things is, as it happens, what the entire series is about.
By day I run product for data strategy and operations at Condé Nast, where the brief is customer identity: the unglamorous business of establishing that the person reading on a phone in Mumbai and the one subscribing on a laptop in London are the same human being. Essentially the ‘slow work’ of turning unknown into known, in various stages. The glamorous parts of my day: developing the single customer view, and harnessing that data to optimise for amplified engagement and revenue across multiple lines and brands.
Twenty-four years of it now, across product, data and technology - client side and agency side, in media and publishing, CPG, insurance, automotive, FMCG and telecom, across North America, Europe and Asia. Enough time in front of CXOs to have learned that a business case travels further than an architecture diagram, and enough time behind them to know the diagram still has to be right.
This project was my evenings. It brings together the triumvirate - my love for the world of finance and markets, my drive to build a production grade system using the latest AI toolset, and the itch to discover first hand what these tools are truly capable of. And the only honest way to find out what these tools can carry is to hand them something that can lose real money.
Questions, disagreements, war stories from your own build, or a conversation about senior product leadership and AI-delivery roles - all welcome.