This is not an AI piece, and it is the most important thing I have published here. It is about the difference between a number you modelled and a number you measured — which is the only question this publication really asks.
Over several months I built a backtest lab and ran 19 systematic strategies across Indian equities, index futures, crypto and NIFTY options. Everything out-of-sample, everything net of costs and, for crypto, net of tax. The point was never to find a money machine. It was to find out cheaply what was true before risking anything.
First finding: the cost floor eats almost everything
Fourteen of the nineteen produced a real gross signal and still lost money. Not because the signal was imaginary — because the instrument carrying it cost more to trade than the signal was worth.
| Instrument | Round-trip cost | Typical gross edge found | Net |
|---|---|---|---|
| Index futures NIFTY / BANKNIFTY | ~6 bps | 1.8 bps | negative |
| Large-cap equity delivery | 14–31 bps | 2.7–5 bps | negative |
| Mid-cap equity | higher + wide spreads | 9.6 bps | unverifiable |
| Crypto India, 1% TDS per trade | ~100 bps | 9.7 bps | negative |
A 1% transaction tax against a 9.7 bps edge is not a headwind, it is a wall. The single crypto strategy that survived out-of-sample became a drawdown reducer rather than a profit source once tax was modelled honestly.
The pattern was consistent and slightly funny once I saw it. Cheap instruments didn't carry enough signal; instruments that carried signal cost too much to trade. Index futures are cheap and nearly signal-free. Mid-caps have real edge and spreads that eat it. Crypto has the strongest raw momentum and a 1% tax per trade.
Every gross edge I found was real. Every one of them sat below the floor of the thing that carried it.
That is also, I think, the honest explanation for the regulator's figure that around 90% of retail derivatives traders lose money. They are not all wrong about direction. They are paying more for access than the access is worth.
Second finding: the one that worked wasn't what I measured
Three strategies did clear costs, and they shared a property: none of them predicts direction. The strongest was the volatility risk premium — selling index options. India VIX averages about 15.2% implied against roughly 13.1% realised, a structural two-point spread that option sellers collect. You are not forecasting the market; you are selling insurance on it.
On my proxy the regime-filtered version scored Sharpe 2.74 out-of-sample, positive on 78% of days. That is an extraordinary number, and extraordinary numbers are where you should get suspicious rather than excited.
So I rebuilt it on real NIFTY option prices — 100 weekly iron condors, January 2024 to November 2025, actual settlement data rather than a variance proxy.
| Measurement | Sharpe | What it was built on |
|---|---|---|
| Variance-premium proxy | 2.74 | Index data, modelled premium |
| Real options, unfiltered | 0.35 | 100 weekly condors, actual prices |
| Real options, regime-filtered | 0.49 | Same, calm-uptrend filter |
Net profit was genuinely positive — about ₹29,700 per lot unfiltered, 79% of weeks winning. The direction of the edge survived. The magnitude did not.
The strategy was still profitable. It was just seven times less good than the model claimed. And the reason matters more than the number: the proxy assumed you capture the whole implied-realised spread, while real options make you pay bid-ask on cheap far-out-of-the-money contracts, post margin, and accept gap risk. Every one of those is invisible in a variance proxy and unavoidable in a trade.
Third finding: the tail is where the money actually lives
The real-data run exposed something the Sharpe number hides. Across 100 weeks, the five worst weeks lost ₹92,000. Strip those five out and the same strategy made ₹122,000. Five weeks out of a hundred were erasing about three quarters of two years of profit.
My instinct was to filter — predict the bad weeks and sit them out. That helped a little and never enough. What actually worked was changing the structure rather than the timing: sell further out of the money, and buy protective wings so the loss is bounded by construction instead of by judgement.
| Structure | Win rate | Net / lot | Worst week | Sharpe |
|---|---|---|---|---|
| 1.5% OTM / 300 wing baseline | 74% | ₹21k | −₹20.7k | 0.17 |
| 2.0% OTM / 500 wing | 85% | ₹169k | −₹21.3k | 1.45 |
| 2.5% OTM / 500 wing tamed | 92% | ₹168k | −₹15.8k | 2.22 |
The tamed structure was positive in every year tested. In the Feb–Jun 2026 window where the baseline lost ₹6,900, it made ₹58,000. Same signal, same market — different construction.
That is the most useful thing I learned in the whole exercise. I had been treating tail risk as a forecasting problem and it was a construction problem. You do not need to know which week explodes if the loss is capped before it starts.
What I'd want stated against my own numbers
Since the entire piece is an argument for distrusting confident backtests, here is the case against this one.
There is no crash in the sample. Free NSE archives only reach back to 2024, so the data contains nothing like 2020 or 2008. A 2.5%-out-of-the-money short option will be breached in a real crash. The wing bounds the loss, but a 92% win rate mathematically guarantees the rare losses are large.
Every tuned parameter is an overfit risk. The regime thresholds, the lookbacks, the specific strike distance — each was chosen partly because it worked on this data.
The realistic number is lower again. Fuller slippage on thin far-OTM contracts plus one genuine crash would likely land a well-run version around Sharpe 1 to 1.5. That is still good. It is not 2.74, and it is emphatically not 3.6, which is what the full combined book scored on proxies before real data intervened.
What transfers beyond trading
- A modelled number and a measured number are different kinds of object. Mine differed by 87% with no error in the logic.
- Gross edge is not edge. Whatever you are measuring, the cost floor of the thing carrying it decides whether anything survives.
- Suspicion should scale with the result. Sharpe 2.74 was not a triumph, it was a signal that the measurement was wrong.
- Bounded downside beats predicted downside. Structure, not forecasting.
- Not transferable: any of these specific numbers. They are one market, one period, one operator's costs.
The same shape shows up in the AI work this site usually covers. Reasoning from a published price sheet is a proxy; reading your own invoices is a measurement. When I finally did the latter, output tokens turned out to be 8% of my bill rather than the dominant cost every optimisation guide implies. Same lesson, different ledger.
Method & disclosure
Nineteen strategies, all coded and re-runnable, tested out-of-sample and net of costs (and of tax for crypto). Price data from free sources — yfinance for equities, futures and crypto; NSE bhavcopy and live chain data for the real options work. The full scorecard including the rejected strategies is kept as an internal truth record.
The proxy figures are variance-premium approximations on index data. The real-options figures come from actual NIFTY option prices: 100 weekly condors (Jan 2024 – Nov 2025) for the first comparison, 132 trades (2024 – mid 2026) for the structure test. Currency is Indian rupees; the market is NSE.
This is not investment advice and none of it is a recommendation. It is a record of what a measurement exercise returned, published mainly because the gap between the modelled and measured numbers is worth knowing about. No positions are being solicited, no product is being sold, and no vendor or broker has any relationship with this publication.