The graveyard

One hundred and eighty-seven ideas tested.
What each one failed at.

Two kinds of stone live here. Some are strategies people sell courses about, rebuilt and tested on their own, after trading costs. Most are changes we tried on our own engine, or things we put beside it, judged only on whether the engine did better with them. A stone of the second kind says nothing about the idea on its own. Every stone says which kind it is, and a few say they were only tested on a small account or in one strict form. The cause of death is listed on the stone, and the corner plate ranks how close it came: S-RANK lived — one, at the bottom · A-RANK near miss — a real effect that died at equal risk or on the mechanism check · B-RANK lost — worse than the engine on the measure that matters · C-RANK could not be run — no contract, no data, no bid. The small word under each plate is how hard it was tested: full battery, validated, measured, reviewed. Hover any plate for the definition. Click a stone to hold it. Survivor at the bottom.

187CLOSED
3ON PAROLE
1ARMED & LIVE

First visit? Start with three: buy the dip, cut leverage on the VIX, a trailing stop instead of the monthly exit. Then the one that lived, at the bottom.

what it failed at
one hundred and eighty-seven stones · newest first · click a stone to link to it
The stones

Two words that come up on every stone. Sharpe is return per unit of risk; a higher number means more return for the same swings, and we treat a difference smaller than 0.03 as nothing. Worst fall is the deepest drop from a high before the book recovered. Every stone shows the verdict first; the full autopsy with the numbers folds underneath.

B-RANKmeasured

The engine selling covered calls on its own shares

The premiums it collected were smaller than what it had to pay back when the market ran.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A covered call means selling someone the right to buy your shares at a higher price, for a fee paid today. Each month the engine sold that right on the Nasdaq shares it already held, at a price 5% above the market, and kept the fee. Over 83 months it collected $153,181 in fees. But when the Nasdaq rose past the price, it owed the difference: $176,361 at the end of those months, plus $39,894 buying the rights back when the engine sold shares. From $250,000 in 2018 it ended with $1,274,664 against $1,407,789 without the calls. The worst fall was barely smaller (-26.8% against -27.5%). The engine makes its money in the strong months, and those are the months a covered call gives away.

R-147 (registered and sealed before it ran) · ULTRA as shipped, real end-of-day bid and ask from our option archive, calls sold at the bid on the day after each monthly expiration, one per 100 shares held · sold 2% above instead: $1,238,651 · lost in both halves · FAIL · scripts/research/r147
B-RANKmeasured

Holding real bitcoin on Robinhood instead of the bitcoin fund

It made more money, but its worst fall came out a hair deeper than the limit we set before testing it.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The bitcoin slice checks once a month whether bitcoin is above its long average. Held in the bitcoin fund, the trade waits for the first weekday of the month and pays the fund’s yearly fee. Robinhood’s agent account can trade real bitcoin any day, so the trade could happen right at the month’s last close, weekends included, at the cost of Robinhood’s markup on every trade. From $10,000 in mid-2015 the bitcoin itself grew to $3,986,358 against $3,417,234 in the fund, and still $3,714,889 if the markup were 1%. But its worst fall was -72.6% against -71.6%, and the rule we wrote before the test allowed one point deeper, no more. It traded only about once a year, so the result rests on a few trades.

R-145 (registered and sealed before it ran) · the frozen sleeve rule, sleeve alone, 0.5% markup per trade (published range 0.35%–1%), fund at 0.25% a year · 14 trades in eleven years · worst fall -72.61% against -71.58% · FAIL · scripts/research/r145
B-RANKmeasured

A long-dated option instead of the two-times fund

It won over 25 years and over the last 12, but lost from 2001 to 2013. It had to win in both halves.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

ULTRA gets its extra push from a fund that holds twice the Nasdaq, which charges a yearly fee and resets every day. Once Robinhood stopped charging for options, we tried getting the same push from a long-dated call option on the Nasdaq fund instead: bought about two years out and swapped for a new one when a year was left. From $100,000 in 2001 it grew to $8,380,192 against $7,029,297, and its worst fall was a little smaller (-26.2% against -27.4%). But from 2001 to 2013 it ended with less: $534,051 against $555,805. We priced the options with a formula, because we do not own a history of real option prices, so a near miss is not enough to buy one. One contract also costs about $29,000, so it could only ever help accounts of $100,000 and up. Retested on real option prices (R-144, 1 October 2026): with every option bought at the asking price and sold at the bid, from $250,000 in mid-2012 the option version grew to $2,563,139 against $5,576,187 with the two-times fund. Crossing the gap between buy and sell prices cost $604,833 along the way: in 2012 and 2014 that gap was about 15% of the option’s price, not the 2% the formula assumed. Real prices settled it.

R-143 (registered and sealed before it ran) · ULTRA as shipped, fractional, bill-rate cash, no commissions; calls priced by Black-Scholes at the Nasdaq volatility index plus 3, 1% of the premium paid on every buy and sell, delta 0.85 at two years, rolled under one year · second stretch 2014–2026 24.49%/yr against 22.29% · first stretch 2001–2013 13.89% against 14.24% · FAIL · scripts/research/r143
B-RANKvalidated

Rebalancing more often once trading is free

With no commissions it could hug its target more closely, and every version fell harder in the worst stretch.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine only adjusts when its holding drifts outside a band around the target, partly to save on trading costs. At a broker with no commissions and fractional shares, that reason looked gone, so we tried a band half as wide and a quarter as wide. Some versions made a little more money: ULTRA with the quarter-width band grew 17.53% a year against 16.94%. But the worst fall got deeper in every case, by 1.1 to 3.2 points (ULTRA: -28.6% against -26.0%). Adjusting more often means chasing the market’s swings, so the band stays as it is. The same test found the engine needs no large account at these costs: from $250 it ran within a fraction of a percent a year of the $25,000 record.

R-141 (registered before it ran; r012 simulate, published settings, $0 commission, fractional shares, 5 bp spread and slippage; $1,000 and $25,000, March 1999 to September 2026, both halves) · SELECT half band 14.82%/yr, -23.2% against 14.81%, -21.6% · FAIL · scripts/research/r141
A-RANKvalidated

Getting back in early, tested on sixty years of the S&P 500 and the Nasdaq Composite

On decades it was never designed on it helped a little almost everywhere, but not by enough to tell apart from luck.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The two graves above had only three bear markets to learn from. This one ran the engine’s rules on the S&P 500 from 1960 and the Nasdaq Composite from 1972, decades none of these rules were designed on, with the early seventies, 1987 and many more fake rallies in them. Going back in early at half size earned a little more in 7 of 8 cases and made the worst fall a little smaller in all 8. But it improved the return for the risk by only about 0.02, under the 0.03 we count as real, everywhere except the S&P 500 before 1999. A nudge in the right direction, too small to change the engine for.

R-139 (registered before it ran; the engine’s timing and sizing on index prices without dividends, not the product) · S&P 500 1960 to 1998, SELECT settings: engine 7.74%/yr, -39.2%; early re-entry 8.24%/yr, -37.5% · Nasdaq Composite 1972 to 2026: 13.57%/yr against 13.58%/yr · passed 2 of 8 cells · FAIL, near miss · scripts/research/r139
B-RANKvalidated

Getting back in early at half size, with a tighter bail-out

The tight bail-out was shaken out by ordinary dips; the looser half-size version was too small a gain to be real.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Two sharper versions of the grave above: go back in at half size until a month closes 5% above the average, and bail the first day the price closes back under the average. The tight trip wire tripped on ordinary dips: in 2023 it bailed on March 13, the day of the bank scare, and missed the rebound anyway, ending with $210,880 against the engine’s $215,742. The closest version, half size with the looser bail, ended with $223,878 but a worst fall of -22.6% against -21.4%, about 0.16% a year more, and in ULTRA it lost in each half of the history taken on its own. Its whole gain rested on five early re-entries in twenty-seven years.

R-138 (registered after R-137, five versions only; SELECT $5,000, March 1999 to September 2026) · engine $215,742, -21.39%; half size + tight trip wire $210,880, -24.86%; half size + 3% bail $223,878, -22.61% · FAIL · scripts/research/r138
B-RANKvalidated

Getting back in early after a crash, with a bail-out if the rally fails

It caught some real rebounds and bought a fake one: a little more money for a deeper worst fall.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine waits for the Nasdaq to close a month 5% above its long-run average before buying back in, so it sat in cash through the first two months of the 2023 rebound. This version went back in as soon as a month closed above the average, and sold again the first day the price fell 3% below it. It came back a month or two sooner in 2003, 2009 and 2023, and it also went back into the fake rally of May 2008 and fell for seven weeks before the bail-out tripped. From $5,000 in March 1999 it ended with $226,210 against the engine’s $215,742, but its worst fall was -27.1% against -21.4%. The pass mark allowed one point deeper, and only 4 of its eight versions ended with more money; every version that checked weekly lost money.

R-137 (r012 simulate, published settings, $5,000, March 1999 to September 2026; registered before it ran) · SELECT: engine $215,742, -21.39%; early re-entry $226,210, -27.08% · ULTRA: engine $394,326, -26.57%; early re-entry $396,294, -32.09% · FAIL · scripts/research/r137
B-RANKmeasured

Foundry: eleven growth-theme funds on a fifth of the account beside ULTRA

It was offered to make the worst fall smaller. Traded day by day the way a buyer’s copy trades it, it did not, and it cost money; the rebuilt history it was chosen on flattered it.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine alone ended with more money in every stretch, at both account sizes. From $10,000 in mid-2007 it grew to $352,805, against $317,198 with a fifth of the account in Foundry (from $25,000: $888,259 against $824,349). Foundry was there to make the worst fall smaller, and it did not: -25.6% against the engine’s -25.4% over the whole stretch, and -24.0% against -21.5% in the years through 2012. It held memory chips, AI, the seven biggest tech companies, neocloud, humanoid robots, weight-loss drugs, space, circuit boards, pipelines, uranium and robotaxis, one eleventh each, each stepping aside below its own 10-month average. It was in the package for two days and nobody ran it. Retried on Robinhood (R-142, 30 September 2026): with no trading fees at all, the engine alone still ended with more: $358,630 against $314,576 from $10,000 ($905,438 against $801,954 from $25,000), and the worst fall was still no smaller. Fees were never the problem.

WORST FALL SMALLER BY, % (LEFT: DEEPER), FROM $10,0002007 TO 2026-0.2%2007 TO 2012-2.5%2013 TO 2026-0.2%
read the autopsy

An earlier test (R-126) put Foundry at a fifth of the account beside ULTRA and found about the same money with a worst fall about two points smaller, after charging back an estimate of what rebuilding young funds from their holdings today flatters. On that result it was offered, beside ULTRA only. This test (R-136, written down before it ran) replayed the shipped package day by day from July 2007: whole shares, commissions, cash that has to settle before it is spent, splits, dividends and the basket’s own pauses. Most of the funds are under two years old, so each traded its real prices from its first day and the rebuilt history before that, charged the same estimate, with synthetic splits keeping each price between $12 and $60 so whole shares cost what they would on a real fund. The pass mark had two parts, both needed at both sizes in every stretch: a smaller worst fall than the engine alone, and more growth than simply keeping a fifth in cash at the same worst fall. It beat cash in most stretches, by +0.49% a year over the whole stretch from $10,000, but it never made the worst fall smaller, so it fails. A harsher charge for the rebuild changed nothing that matters. What it teaches: a result measured on a rebuilt history and a simplified account can vanish when the real machinery trades it, the same lesson Rotation taught a day earlier; the replay is now the gate for anything offered.

R-136 (R-132 v4 replay machinery, package 2026.10.05; DRAM CHAT MAGS NCLD HUMN OZEM MARS CCML MLPX UX CABZ at 1/11, 10-month exit each, 20% beside ULTRA; real prices from each fund’s first day, R-125 rebuild before it less 4.03%/yr; 2007-07-02 to 2026-09-25; dividends 45 days after the ex-date where undocumented) · from $10,000: engine $352,805, 20.35%/yr, -25.4%; Foundry $317,198, 19.69%/yr, -25.6%; 20% cash $203,625, 16.96%/yr, -22.6% · from $25,000: engine $888,259; Foundry $824,349 · Foundry minus 20% cash at the same worst fall, $10,000: +0.49 / -0.59 / +0.89%/yr (2007 to 2026 / 2007 to 2012 / 2013 to 2026) · FAIL · docs/research/R-136
B-RANKmeasured

Tech and infrastructure: chips, technology, industrials and health care at a tenth of the account

A tenth of the account in the market’s own sectors cannot add money beside the engine, and its small edges were no bigger than the engine’s own path noise; the same lesson as grave 156.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine alone ended with more money in every stretch, at both account sizes. From $10,000 in mid-2007 it grew to $352,805, against $329,513 with a tenth of the account in this basket, $319,527 with a tenth in a broad US market fund, and $306,970 with a tenth in cash (from $25,000: $888,259 against $818,973, $782,118 and $759,211). The basket held a chip fund, a technology fund, an industrials fund and a health care fund at a quarter each, fixed, with no exit. Of the three ways to hold a tenth outside the engine it ended with the most money over the whole stretch, but the rule set before the run asked it to beat both of the others at the same worst fall in every stretch at both sizes, and from $10,000 it lost to both from 2007 to 2012: by 1.75% a year against cash and 0.52% a year against the broad fund.

AHEAD OF 10% IN CASH, % A YEAR, FROM $10,0002007 TO 2026+0.50%2007 TO 2012-1.75%2013 TO 2026+1.03%
read the autopsy

An outside reviewer proposed this basket, and we froze it before any data was loaded (R-133, commit 838e5a61): SOXX, VGT, XLI and XLV at 25% each, no exit, at 10% of the account beside ULTRA, run as a buyer’s own basket by the real package, day by day from July 2007, with whole shares, commissions, cash that has to settle before it is spent, splits and dividends. One declared change: the reviewer asked for a quarterly rebalance, and the product checks monthly and trades only when a holding drifts more than a quarter from its target, so it was tested the way a buyer would run it. It was compared with the engine alone, the engine on 90% with 10% in cash earning the Treasury bill rate, and the engine on 90% with 10% in two total US market funds. The pass mark: beat both 90% versions at the same worst fall, from 2007 to 2026 and in each half, from $10,000 and from $25,000. It passed every stretch from $25,000 and failed from $10,000 in 2007 to 2012. Reported without letting it decide: the engine itself makes slightly different calls at different account sizes, because whole shares and its rebalancing thresholds land differently, and in 2015 and 2016 that alone moved the worst fall by about 1.6 points between two books with the same rule; the basket’s edges were of the same size. Two assumptions, stated: dividends were paid on their documented dates where known and 45 days after the ex-date otherwise, and cash earned the bill rate at no cost, which a small real account does not. This retires this basket under these rules; it does not show that every sector basket must fail.

R-133 (R-132 v4 replay of package 2026.09.29; SOXX 25 / VGT 25 / XLI 25 / XLV 25, no exit, 10%, monthly check with the 25% band; 2007-07-02 to 2026-09-25) · from $10,000: engine $352,805, 20.35%/yr, -25.4%; basket $329,513, 19.93%/yr, -26.9%; broad $319,527, 19.74%/yr; cash $306,970, 19.49%/yr · from $25,000: engine $888,259; basket $818,973; broad $782,118; cash $759,211 · basket minus cash at the same worst fall, $10,000: +0.50 / -1.75 / +1.03%/yr; minus broad: +0.19 / -0.52 / +0.40%/yr (2007 to 2026 / 2007 to 2012 / 2013 to 2026) · FAIL · docs/research/R-133
B-RANKmeasured

Four Corners: chips, health, energy and staples beside the engine

A basket growing at half the engine’s pace cannot add money beside it, and from 2000 to 2012 simply holding less of the engine lowered risk more cheaply; the same lesson as grave 156.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine alone ended with more money in every stretch we measured: 16.06% a year from mid-2000 to 2026, against 13.78% with 30% of the account in Four Corners beside it (9.19% against 6.73% from 2000 to 2012, 22.16% against 20.18% from 2013 to 2026). Four Corners was a second version of the Foundry: 40% in a chip fund and 20% each in health care, energy and consumer staples, each holding with its own exit, all real funds with no stand-ins. Beside the engine it had to beat simply holding less of the engine, which falls just as far at worst, over the whole stretch and in both halves. It came out ahead from 2013 on, by 2.35% a year, and behind from 2000 to 2012, by 1.45% a year. On its own it grew 7.72% a year, about half the engine’s pace, with a worst fall of 41.8%.

AHEAD OF HOLDING LESS ENGINE, % A YEAR2000 TO 2026+0.60%2000 TO 2012-1.45%2013 TO 2026+2.35%
read the autopsy

We wrote the rules and the pass mark down before any data was loaded, so nothing could be tuned afterwards (R-112, commit 2907b5ce). One basket was fixed in advance, with no search: SMH 40, XLV 20, XLE 20, XLP 20, rebalanced monthly, each holding stepping aside while its month-end is at or below its 10-month average, exactly as the Foundry preset runs. Only the real funds were used. The health, energy and staples funds have prices from January 1999; the chip fund lists in June 2000, so the test starts in July 2000. The pass mark, the same as the Foundry’s own history test (grave 177): the engine with 30% in the basket had to earn more a year than the engine turned down to the same worst fall, the rest in Treasury bills, from mid-2000 to 2026 and in each half, 2000 to 2012 and 2013 to 2026. It won over the whole stretch (13.78% a year against 13.18% for 78% of the engine) and from 2013 (20.18% against 17.83%), and lost from 2000 to 2012 (6.73% against 8.18% for 83% of the engine). One lost half is a fail. We said before the run that we expected a fail: a sector all-weather book with health, staples, utilities and energy had already failed beside this engine (grave 156). Reported without letting them decide: in the tech bust, July 2000 to October 2002, the engine alone lost 1.7%, the engine with Four Corners 13.9%, and Four Corners on its own 38.0%. In the 2008 crash the mix lost 14.9% against 15.9% for the engine alone, and in 2022 14.0% against 19.5%. Equal quarters in each fund also lost from 2000 to 2012, by 0.92% a year; with the exits turned off the basket was behind over the whole stretch as well, by 1.32% a year.

R-112 (R-012 harness, the engine with cash model v2; SMH 40 / XLV 20 / XLE 20 / XLP 20, per-holding 10-month exit, monthly rebalance; real funds only, SMH’s weight shared among the other three before its first price; 2000-07-03 to 2026-08-31) · 2000 to 2026: engine 16.06%/yr, -27.5%; v2 7.72%/yr, -41.8%; 70/30 13.78%/yr, -22.0%; dial 78% engine 13.18%/yr, -21.9%; 70/30 minus dial +0.60%/yr · 2000 to 2012: engine 9.19%/yr; v2 0.36%/yr; 70/30 6.73%/yr, -22.0%; dial 83% 8.18%/yr; 70/30 minus dial -1.45%/yr · 2013 to 2026: engine 22.16%/yr; v2 14.68%/yr; 70/30 20.18%/yr, -22.0%; dial 78% 17.83%/yr; 70/30 minus dial +2.35%/yr · reported, 2000-07 to 2002-10: engine -1.7%, 70/30 -13.9%, v2 -38.0% · 2007-10 to 2009-03: engine -15.9%, 70/30 -14.9% · 2022-01 to 2022-10: engine -19.5%, 70/30 -14.0% · reported, equal 25s minus dial: +0.41 / -0.92 / +1.62%/yr; exit off minus dial: -1.32 / -1.95 / +0.60%/yr · FAIL · docs/research/R-112
B-RANKmeasured

The Foundry through the tech bust, beside the engine

The theme paid in the AI boom and cost in the tech bust; beside the engine it is a bet on the theme, not an edge.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

From 2000 to 2026 the engine alone ended with more money: 15.55% a year, against 14.49% with 30% of the account in the Foundry beside it. The Foundry is our AI build-out basket: chip makers, the power grid, Alphabet, a network-gear maker, uranium, space, quantum computing and robotics, each holding with its own exit. Its funds are young, so we rebuilt it on older stand-in funds back to 2000, to see how it would have done in a tech bust. Beside the engine it had to beat simply holding less of the engine, which falls just as far at worst, over the whole stretch and in both halves. It came out ahead from 2013 on, by 1.04% a year, and behind from 2000 to 2012, by 0.37% a year. In the tech bust, March 2000 to October 2002, the engine alone lost 8.7%, the engine with the Foundry 15.4%, and the Foundry on its own 29.8%.

AHEAD OF HOLDING LESS ENGINE, % A YEAR2000 TO 2026+0.30%2000 TO 2012-0.37%2013 TO 2026+1.04%
read the autopsy

We wrote the rules and the pass mark down before any data was loaded, so nothing could be tuned afterwards (R-110, commit 6b735ec1). The basket is the Foundry exactly as set: SMH 35, GRID 25, GOOGL 15, ANET 10, URA 5, SPCX 4, QTUM 3, BOTZ 3, rebalanced monthly, each holding stepping aside when it closes a month at or below its 10-month average. Until each real fund has a price, an older one stands in: the Philadelphia chip index for SMH (price only, so its dividends are missing), half utilities and half industrials for GRID, the Nasdaq-100 for Alphabet before 2004 and for the space, quantum and robotics funds, Cisco for Arista before 2014, and Cameco for URA before 2010. The stand-ins were picked by someone who knew how 2000 to 2002 went; we say so rather than hide it. The pass mark: the engine with 30% in the Foundry had to earn more a year than the engine turned down to the same worst fall (89% of the engine, the rest in Treasury bills) over 2000 to 2026 and in each half, 2000 to 2012 and 2013 to 2026. It won over the whole stretch (14.49% a year against 14.19%) and from 2013 (21.05% against 20.01%), and lost from 2000 to 2012 (7.49% against 7.86%). One lost half is a fail. The engine alone beat both: 15.55% a year with a worst fall of −27.6%, against the mix’s −24.9%. Reported without letting them decide: the Foundry’s own exits are what kept it survivable. In the tech bust the Foundry alone lost 29.8% with its exits and 61.2% without them. Without the exits it would have made more over the whole stretch (13.34% a year against 11.30%), but fallen 72.0% at worst against 37.3%. In the 2008 crash the mix lost 15.4% against 15.9% for the engine alone, and in 2022 18.7% against 19.5%, so it cushioned those two falls a little and deepened the tech bust. Smaller slices were measured too and are listed below; choosing a size after seeing them would be tuning, so none of them counts. This test changes nothing live: the Foundry’s sealed forward watch from 2026-09-24 keeps running.

R-110 (R-012 harness, the engine with cash model v2; Foundry SMH 35 / GRID 25 / GOOGL 15 / ANET 10 / URA 5 / SPCX 4 / QTUM 3 / BOTZ 3, per-holding 10-month exit, monthly rebalance; stand-ins until each fund’s first price: ^SOX for SMH, XLU/XLI 50/50 for GRID, QQQ for GOOGL, SPCX, QTUM and BOTZ, CSCO for ANET, CCJ for URA; 2000-01-03 to 2026-08-31) · 2000 to 2026: engine 15.55%/yr, -27.6%; Foundry 11.30%/yr, -37.3%; 70/30 14.49%/yr, -24.9%; dial 89% engine 14.19%/yr, -24.8%; 70/30 minus dial +0.30%/yr · 2000 to 2012: engine 8.40%/yr; Foundry 4.73%/yr; 70/30 7.49%/yr, -23.3%; dial 89% 7.86%/yr; 70/30 minus dial -0.37%/yr · 2013 to 2026: engine 22.16%/yr; Foundry 17.71%/yr; 70/30 21.05%/yr, -24.9%; dial 89% 20.01%/yr; 70/30 minus dial +1.04%/yr · reported, 2000-03 to 2002-10: engine -8.7%, 70/30 -15.4%, Foundry -29.8%, Foundry exit off -61.2% (worst -72.0%) · 2007-10 to 2009-03: engine -15.9%, 70/30 -15.4% · 2022-01 to 2022-10: engine -19.5%, 70/30 -18.7% · reported, sizes 2000 to 2026 minus dial: 10% +0.52%/yr, 20% +0.79%/yr, 30% +0.30%/yr · Foundry alone exit off 2000 to 2026 13.34%/yr, -72.0% · FAIL · docs/research/R-110
B-RANKmeasured

Copying what members of Congress buy, beside the engine

Neither fund beat simply holding less of the engine, and one of them closed.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Two funds were built to buy what members of Congress disclose buying: NANC follows Democrats’ trades, KRUZ Republicans’. We put 20% of the account in each, and in a half-and-half mix, beside the engine, and compared that with holding a little less of the engine, which falls just as far at worst. NANC came out 0.29% a year ahead over the whole window, March 2023 to August 2026, but behind since 2025 (−0.21% a year). KRUZ closed in March 2025; in the stretch it can be judged on, to the end of 2024, it and the mix both came out behind. On its own NANC made 25.04% a year, more than the S&P 500’s 22.45% and less than the Nasdaq-100’s 30.08%.

AHEAD OF HOLDING LESS ENGINE, % A YEARNANC, WHOLE WINDOW+0.29%NANC, TO END 2024+0.20%NANC, SINCE 2025-0.21%KRUZ, TO END 2024-0.36%50/50 MIX, TO END 2024-0.08%
read the autopsy

We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards (R-104, commit f21233f4; amended before any return was computed, 602e971b, to price KRUZ from Tiingo, because Yahoo has no history for it). Both funds listed in February 2023, so the window is short: March 2023 to August 2026, in two stretches, to the end of 2024 and since 2025. Each fund was held with no exit, 20% beside the engine, re-sized monthly. To pass, a seat had to earn more a year than the engine turned down to the same worst fall (the rest in Treasury bills) over the whole window and both stretches, and the fund alone had to beat the S&P 500 over the whole window. A pass would only have earned a forward watch. NANC beat the turned-down engine over the whole window (32.69% a year against 32.40% for 93% of the engine) and to the end of 2024 (0.20% ahead), then trailed since 2025, 21.56% against 21.77% for 99% of the engine. KRUZ’s last close was March 20, 2025, so its record and the mix’s stop there, and any figure for them that runs past it would compare a shortened record with the engine’s full one. We quote them only to the end of 2024. There, KRUZ beside the engine made 40.74% a year against 41.10% for 86% of the engine, and the mix 42.36% against 42.44%. On its own KRUZ made 15.48% a year in that stretch, while the Nasdaq-100 made 36.80%.

R-104 (R-012 harness, the engine; NANC from Yahoo, KRUZ from Tiingo; adjusted closes; seat held, no exit, 20%, monthly re-size; 2023-03-01 to 2026-08-31; KRUZ last close 2025-03-20, so KRUZ and NANC+KRUZ are quoted for 2023-03 to 2024-12 only) · NANC whole window: alone 25.04%/yr, -20.9%; 80/20 32.69%/yr, -26.1%; dial 93% 32.40%/yr; SPY 22.45%/yr; QQQ 30.08%/yr; engine 34.51%/yr, -27.7% · NANC to 2024-12: 80/20 43.99%/yr vs dial 92% 43.79%/yr · NANC from 2025-01: 80/20 21.56%/yr vs dial 99% 21.77%/yr · KRUZ to 2024-12: alone 15.48%/yr, -10.0%; 80/20 40.74%/yr vs dial 86% 41.10%/yr · NANC+KRUZ to 2024-12: alone 22.82%/yr; 80/20 42.36%/yr vs dial 89% 42.44%/yr · QQQ to 2024-12 36.80%/yr · all three FAIL · docs/research/R-104
B-RANKmeasured

Ether in the bitcoin seat, or sharing it

Ether won the 2018–2021 boom, lost since, and adds bitcoin’s risk, not a new one.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The live account holds bitcoin beside the engine, and only while its price is above its 200-day average. We gave ether exactly the same rule and tried it two ways: in bitcoin’s place, and splitting the seat half and half. From November 2018 to the end of 2021 the book with ether made 54.05% a year against 39.41% with bitcoin. Since 2022 it made 15.57% against 20.32%, and less than simply holding a little less of the engine (16.92%). The two coins move together most days (a correlation of 0.82, where 1 is lockstep), and sharing the seat made the book’s worst fall deeper, not shallower: −29.5% against −28.2% with bitcoin alone.

ENGINE 80 + 20 IN THE SEAT, RETURN A YEARWITH BITCOIN, TO 202139.4%WITH ETHER, TO 202154.0%WITH BITCOIN, SINCE 202220.3%WITH ETHER, SINCE 202215.6%
read the autopsy

We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards (R-103, commit 2172c607). The engine is the same full-history harness as every other seat test; each coin is held only while its close on the first session of the month is above its 200-day average, otherwise the seat earns the Treasury-bill rate; the book is re-sized monthly and each switch costs 0.10%. The window starts in November 2018, the first month with a full year of ether prices for the average, and runs to August 2026, in two stretches: to December 2021, and since 2022. To pass, ether had to do two things in the whole window and in both stretches: earn more a year than the engine turned down to the same worst fall (the rest in Treasury bills), and beat the bitcoin seat on return divided by worst fall. In bitcoin’s place it did both in the first stretch (1.759 against 1.398), then failed both since 2022: 0.619 against 0.779, and 15.57% a year against 16.92% for 91% of the engine alone. Over the whole window it came to 0.970 against bitcoin’s 0.979. Sharing the seat 10/10 tied bitcoin over the whole window (0.979 against 0.979), and a tie is not a pass; since 2022 it trailed, 0.702 against 0.779. The bitcoin seat itself is unchanged by this test.

R-103 (R-012 harness, the engine with cash model v2; Yahoo BTC-USD and ETH-USD at the engine’s sessions; first-session close above the 200-day average, else 3-month T-bill; monthly re-size; 10bp per flip; 2018-11-01 to 2026-08-31) · whole window: engine alone 22.72%/yr, -27.7%; 80 / BTC 20 27.60%/yr, -28.2%, ratio 0.979, dial 100% 22.72%/yr; 80 / ETH 20 29.81%/yr, -30.7%, ratio 0.970, dial 100% 22.72%/yr; 80 / BTC 10 / ETH 10 28.86%/yr, -29.5%, ratio 0.979, dial 100% 22.72%/yr · first stretch, to 2021-12: engine alone 30.24%/yr, -23.5%; 80 / BTC 20 39.41%/yr, -28.2%, ratio 1.398, dial 100% 30.24%/yr; 80 / ETH 20 54.05%/yr, -30.7%, ratio 1.759, dial 100% 30.24%/yr; 80 / BTC 10 / ETH 10 46.88%/yr, -29.5%, ratio 1.590, dial 100% 30.24%/yr · second stretch, from 2022-01: engine alone 18.03%/yr, -27.4%; 80 / BTC 20 20.32%/yr, -26.1%, ratio 0.779, dial 95% 17.42%/yr; 80 / ETH 20 15.57%/yr, -25.2%, ratio 0.619, dial 91% 16.92%/yr; 80 / BTC 10 / ETH 10 17.98%/yr, -25.6%, ratio 0.702, dial 93% 17.17%/yr · BTC/ETH daily-return correlation from 2018-11: 0.82 · replace: FAIL; join 10/10: FAIL · docs/research/R-103
B-RANKmeasured

Buying after a big earnings beat and holding for three months

The drift came from counting only today’s winners; checked fairly, it was zero.
failed on its ownTested as a strategy on its own, after trading costs.

When a company beats its earnings forecast by a wide margin, its stock is said to keep drifting up for weeks. We bought at the close after the biggest beats each quarter and held 60 trading days, against simply holding the S&P 500. Tested on the companies in the S&P 500 today, it passed: +2.39% a trade. But today’s members are the ones that did well enough to still be there. From 2010, counting only companies that were in the index at the time and judging each stock against its own usual quarter, the same trade made +0.04%.

GAIN PER TRADE VS THE S&P 500, FROM 2010TODAY’S MEMBERS+1.51%MEMBERS AT THE TIME+0.46%EACH STOCK VS ITSELF+0.58%BOTH CHECKS+0.04%
read the autopsy

Two tests, each with its rules and pass mark written down before it ran, so nothing could be tuned afterwards (R-099, commit dd0a0915; R-101, commit ff399e80). The first used the 503 companies in the S&P 500 today and their reports from 2003: buy at the close of the first session that could react to the report, sell 60 sessions later, 0.10% in costs. It passed every registered condition: +2.39% a trade over 8,172 trades, a t of 6.2 against a bar of 3 (t measures how far a result stands above its own noise; each quarter counts once so overlapping trades cannot inflate it), +3.15% in the first stretch (2003–2014) and +1.60% in the second (2015–2026), and +1.74% since 2021. It said at the time that a list of today’s survivors flatters any buy signal, and one number showed it: even the fifth of reports with the worst surprises beat the S&P 500, by 0.91%. The second test ran two registered checks on the same trades from 2010. Counting only companies that were members on the day (the list rebuilt from the index’s changes table) cut the gain to +0.46% (t 1.3). Judging each stock against its own average quarter, which removes the lift a survivor gets every quarter, cut it to +0.58% (t 1.5). With both, the pass mark was a t above 2.5 and a gain in each stretch; it made +0.04% (t 0.1), +0.37% in the first stretch (2010–2017) and −0.27% in the second (2018–2026). The registration said a fail here ends the idea, so the planned test of a real ten-stock sleeve beside the engine was not run. Data: Tiingo, which prices companies that have left the index, has earnings data only for the 30 Dow stocks, so a full rebuild from the companies’ own SEC filings was left for a later test; we did not run it, because the checks we could run already brought the result to zero.

R-099 (today’s 503 S&P 500 members, Yahoo earnings dates and surprise %, reports 2003-01-01 to 2026-05-31, top surprise quintile within each calendar quarter, day-0 close to day +60, abnormal vs SPY after 10bp, t over quarters): top +2.39%, t 6.23, 94 quarters, 8,172 trades; first stretch 2003–2014 +3.15% (t 5.71); second stretch 2015–2026 +1.60% (t 3.12); since 2021 +1.74% (t 2.43); bottom quintile +0.91% (t 2.50) · R-101 (same frozen inputs, from 2010; 30,381 reports, 24,636 by members on the day): no check +1.51% (t 3.82, 6,089 trades), first stretch 2010–2017 +1.32%, second stretch 2018–2026 +1.69%; members on the day +0.46% (t 1.25, 4,940 trades), first stretch 2010–2017 +0.77%, second stretch 2018–2026 +0.18%; each stock vs itself +0.58% (t 1.51, 6,089 trades), first stretch 2010–2017 +0.60%, second stretch 2018–2026 +0.56%; both (gating) +0.04% (t 0.11, 4,940 trades), first stretch 2010–2017 +0.37%, second stretch 2018–2026 -0.27% · Part 2 (a 10-name sleeve at $25,000 and $5,000 beside the engine) not run, as registered · docs/research/R-099, R-101
B-RANKmeasured

The Globe: our own world fund, a bit more aggressive, beside the engine

Beside the engine it did no better than simply holding less of the engine.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The idea was our own version of a whole-world fund, tilted to be a bit more aggressive: US growth, US small value, developed and emerging markets abroad, international small value and momentum, each with its own exit. It was meant as insurance for a decade when US tech goes nowhere. We put 30% of the account in it beside the engine and compared that with holding 85% of the engine and the rest in Treasury bills, which falls just as far at worst. The plain cut-back made 0.17% a year more, and it won in both halves of the test as well.

RETURN A YEAR · WORST FALL, 2003–2026ENGINE ALONE18.2% · −28%85% ENGINE, REST IN BILLS15.9% · −24%ENGINE 70 + GLOBE 3015.7% · −24%GLOBE ALONE9.3% · −21%
read the autopsy

We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards (R-100, commit a1263e16). The basket: 25% Nasdaq-100, 15% US small value, 20% developed markets, 10% international small caps, 20% emerging markets, 10% momentum, rebalanced monthly, each holding stepping aside when it closes a month below its 10-month average. The funds we would actually buy are young, so the test used older funds that track the same things (QQQ, IJS, EFA, SCZ, EEM, MTUM), from May 2003, when the emerging-markets fund has its first full month. Pass mark: the engine with 30% in the Globe had to earn more a year than the engine turned down to the same worst fall, over the whole window and in each stretch of it, 2003–2014 and 2015–2026. It trailed by 0.35% a year in the first stretch (against 90% engine) and 0.03% in the second (against 81%). Two things we reported without letting them decide. On its own since July 2008 the Globe was not a more aggressive world fund but a calmer one: 7.8% a year with a worst fall of −21.0%, against 9.0% and −49.8% for the whole-world fund VT. The exits bought the calm and cost the return. And the real funds, which exist only since October 2019, were too young for their 10-month exits to work in the 2020 crash; they rode it down, a worst fall of −32.4% against −17.3% for the stand-ins. A basket of stocks falls with the engine in a crash, so its only chance was a stretch when markets abroad lead; the first stretch (2003 to 2014) holds such years, and over that whole stretch it still trailed. Same verdict as the chips-and-energy pair beside the engine (grave 169).

R-100 (R-012 harness, the engine with cash model v2; stand-ins QQQ 25 / IJS 15 / EFA 20 / SCZ 10 / EEM 20 / MTUM 10, SCZ from 2007-12 and MTUM from 2013-04 shared pro rata before; per-holding 10-month exit, monthly rebalance; out of the market 0%; 2003-05-01 to 2026-08-31) · whole 2003–2026: engine 18.2%/yr, -27.5%; Globe 9.3%/yr, -21.0%; 70/30 15.71%/yr, -23.8%; dial 85% engine 15.88%/yr, -23.7%; 70/30 minus dial -0.17%/yr; 80/20 16.6%/yr, -24.6% · first stretch 2003–2014: engine 16.6%/yr, -26.1%; Globe 10.2%/yr, -21.0%; 70/30 14.86%/yr, -23.8%; dial 90% engine 15.21%/yr, -23.6%; 70/30 minus dial -0.35%/yr; 80/20 15.5%/yr, -24.6% · second stretch 2015–2026: engine 19.7%/yr, -27.5%; Globe 8.6%/yr, -17.3%; 70/30 16.57%/yr, -22.6%; dial 81% engine 16.60%/yr, -22.6%; 70/30 minus dial -0.03%/yr; 80/20 17.6%/yr, -24.2% · reported, from 2008-07: Globe 7.83%/yr, -21.0%; VT 9.00%/yr, -49.8%; ACWI 8.81%/yr, -51.2% · reported, tradable QQQM/AVUV/VEA/AVDV/AVEM/MTUM from 2019-10: Globe 12.52%/yr, -32.4% (stand-ins 11.82%/yr, -17.3%); 70/30 minus dial (97% engine) -3.07%/yr · docs/research/R-100
B-RANKmeasured

Buying a stock between its S&P 500 announcement and the day index funds buy it

The jump is already in the price before a daily account can buy.
failed on its ownTested as a strategy on its own, after trading costs.

When S&P announces that a company will join the S&P 500, every index fund has to buy it on a set day a few days later. So we tried buying at the first close after the announcement and selling at the close just before the funds buy. Across 208 additions from 2010 to May 2026, held a typical 3 trading days, it lost 0.42% a trade against simply holding the S&P 500 after costs, and only 42% of trades came out ahead. The stock jumps in the first trade after S&P’s evening announcement; by the next close, the earliest a daily account can buy, there is nothing left to catch.

AVERAGE RESULT PER TRADE VS THE S&P 500ALL 208 ADDITIONS-0.42%ADDED 2010–2017-1.27%ADDED 2018–2026+0.12%ALL, AT 30BP COSTS-0.62%
read the autopsy

We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards (R-098, commit bb2e6f36). It grew out of an earlier test that is not a stone of its own. That one (R-097, commit e2f57a7a) asked whether new members keep rising after the day they join. It was void by its own rule: we could price only 79.3% of the 328 additions against a bar of 80%. What it did show was the wrong way round: −1.16% against the S&P 500 twenty trading days after joining. Its rough five sessions before the joining day showed +1.07%, and this test asked whether an ordinary account can catch that. The announcement date is the one printed in S&P’s own press release; announcements come after the close, so the buy is the next close. 208 of 328 additions (63%, bar 60%) had a usable date, prices and at least one session between buy and sell. Costs: 0.10% a round trip, and a harsh 0.30%. Pass mark: an average gain after costs across all trades, strong enough that luck is an unlikely explanation, plus a gain in each of the two stretches of the window and at the harsh cost. The first stretch, additions from 2010–2017, lost 1.27% a trade, and so consistently that chance cannot explain it. The second stretch, 2018–2026, made +0.12%, too little to tell from nothing. An older stone (grave 46) found the classic index-inclusion pop has shrunk from its textbook size; this is the part of it a daily account could still reach, and it is gone.

R-098 (additions to the S&P 500 effective 2010-01-01 to 2026-05-31 from the frozen Wikipedia changes table; announcement dates from the cited S&P DJI press releases, 1–30 days before effective; Yahoo adjusted closes; abnormal = stock minus SPY over the same closes) · usable 208 of 328 (63.4%, bar 60%) · median hold 3 sessions · after 10bp: mean -0.42%, median -0.44%, t -1.40, 42% positive · first stretch 2010–2017 (n 81): -1.27%, t -4.44 · second stretch 2018–2026 (n 127): +0.12%, t +0.26 · after 30bp: -0.62%, t -2.06 · R-097 (void, coverage 79.3% < 80%): day 20 -1.16% (t -2.12, n 260); pre-effective 5 sessions +1.07% (t +2.65) · docs/research/R-097, R-098
B-RANKmeasured

Holding the engine’s fund only overnight

Trading in and out twice a day cost more than the overnight gain.
failed on its ownTested as a strategy on its own, after trading costs.

Much of the Nasdaq’s long-run gain has arrived overnight, between one day’s close and the next morning’s open. So we tried holding the engine’s own fund only overnight: buy at the closing bell, sell at the opening bell, and only in the months the engine’s exit says to be in. Before costs it kept most of the return, 18.7% a year against 19.9% for simply holding it. After paying to trade twice a day it made 12.0% a year, and 3.4% if each fill came in slightly worse. Its worst fall did not improve (−51.9% against −51.7%): the exit already steps aside before the big falls that sitting out the day would dodge.

RETURN A YEAR, 2007–2026HELD, WITH THE EXIT19.9%OVERNIGHT, NO COSTS18.7%OVERNIGHT, AFTER COSTS12.0%SLIGHTLY WORSE FILLS3.4%
read the autopsy

We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards (R-096, commit fa55a3a6). The fund is QLD, the two-times Nasdaq fund the engine holds, on daily prices from 2007 to August 2026. The exit is the engine’s, simplified: be in only in months when the previous month-end Nasdaq close was above its 10-month average. Costs: 0.01% of slippage on each side of each trade at the closing and opening auctions, plus the broker’s per-share commission, at a $25,000 account. Pass mark: return divided by the worst fall had to beat simply holding the fund under the same exit, over the whole window and in each ten-year stretch of it, 2007–2016 and 2017–2026. It passed in the first stretch only, where it made 7.2% a year against 11.2% but fell at worst −30.0% against −51.6%. In the second stretch it made 17.1% against 29.5% with the same worst fall. The overnight effect itself is real: without the exit, overnight-only cut the fund’s worst fall from −83.1% to −52.3%. The exit does that job already (−51.7%) and keeps the daytime return, so the two cannot be stacked. Taxes were not counted; in a taxable account every overnight gain would be short-term. We buried the same idea on index futures earlier (grave 104); on the fund it fails for the same reason.

R-096 (Yahoo daily QQQ/QLD open and close, dividend-adjusted, 2007-01-03 to 2026-08-31; gate: prior month-end QQQ above its 10-month average; IBKR tiered $0.0035/share, $0.35 minimum; out of the market earns 0%) · gated QLD held 19.9%/yr, worst -51.7% · gated overnight, no costs 18.7%/yr, worst -51.5% · 1bp/side $25,000 12.0%/yr, worst -51.9% · 1bp/side $5,000 11.8%/yr · 3bp/side $25,000 3.4%/yr, worst -54.0% · return / |worst fall|: held 0.385 vs 0.232 (fail); first stretch 2007–2016 0.217 vs 0.241 (pass); second stretch 2017–2026 0.571 vs 0.330 (fail) · docs/research/R-096
B-RANKmeasured

Income, dividend and theme funds as the whole Roth

None of 14 ended with as much as the plain index it comes from, and riding 21 theme booms with an exit never beat simply holding QQQ.
failed on its ownTested as a strategy on its own, after trading costs.

Two questions a Roth owner actually asks. First: skip the engine, hold a dividend fund, a covered-call fund or a robotics fund instead. 14 of them had enough history to judge. Every one ended with less than the index it comes from. QYLD turned $32,000 into $72,744 while QQQ made $205,324. The dividend funds did soften some bad years. They still finished behind. Second: never mind holding a theme, ride it. Step in while it trends up, step out when it breaks, across 21 themes from solar and cannabis to chips and uranium. The exit cut the worst fall in every one of them. It still made less than holding QQQ in every one, by a median 9% a year. The reason is plain once you see it: every boom’s winners end up in the index. A theme fund keeps the losers and the hype as well.

ENDED WITH THIS SHARE OF THE INDEX’S MONEY100% = kept upROBO34%QYLD35%PBP40%BOTZ47%XYLD54%INCOME MIX64%DVY66%SDY72%NOBL72%ARKQ75%SCHD81%DGRO87%VIG88%VTI+THEMES96%
read the autopsy

Registered before each run (R-087, commit 25cb0bae; R-088, commit 8709270d). R-087 held each fund as the entire $32,000 Roth, bought once, payouts reinvested, and asked for either more money than its index or a clearly calmer ride: a better return for the risk and a worst fall at least 5 percentage points shallower (say −30% instead of −35%), across the whole window and both halves of it. No fund met either test in any window. 5 were too young to judge and were reported without a verdict: JEPI, SPYI, and the new memory and neocloud funds DRAM, DISK and NCLD. R-088 answered the objection that young funds cannot be tested: it tested the pattern instead of the fund, the same monthly trend rule the engine uses, on 21 themes, winners and flops, each against QQQ bought on the same day. A second version stepped in only when the theme was also beating QQQ; it did worse, a median 13% a year behind. The result held in the older themes, the younger ones, and with the three best removed. The same rule applied to QQQ itself is what works, and it is what the engine already does. Looking through to single companies, the rule did beat holding Micron and the chip index; that is one lead, picked by us, and it is being watched forward, not claimed.

R-087 ($32,000, whole shares, $1 per order + 0.05% a side, payouts reinvested, gate window to 2024-12-31, Yahoo adjusted closes frozen with sha256) · 80 VTI + 20 themes from 2016: $95,451 vs SPY $99,660 · ARKQ from 2014: $136,098 vs QQQ $182,423 · BOTZ from 2016: $69,285 vs QQQ $146,013 · DGRO from 2014: $99,059 vs SPY $114,228 · DVY from 2003: $172,258 vs SPY $259,726 · NOBL from 2013: $93,832 vs SPY $129,518 · PBP from 2008: $70,476 vs SPY $178,074 · QYLD from 2014: $72,744 vs QQQ $205,324 · ROBO from 2013: $73,209 vs QQQ $217,206 · SCHD from 2011: $158,817 vs SPY $195,569 · SDY from 2005: $153,061 vs SPY $212,643 · VIG from 2006: $183,024 vs SPY $207,342 · VIG+DVY+PBP from 2008: $113,661 vs SPY $178,074 · XYLD from 2013: $76,704 vs SPY $142,309 · 0 of 14 · The same test at $10,000 (re-run 2026-09-25, scripts/research/r087/results_10k.json; whole shares and $1 orders, so not a scaled copy): VIG from 2006: $57,121 vs SPY $64,440 · SDY from 2005: $47,714 vs SPY $65,913 · DVY from 2003: $53,716 vs SPY $80,541 · SCHD from 2011: $49,626 vs SPY $60,568 · DGRO from 2014: $30,959 vs SPY $35,378 · NOBL from 2013: $29,288 vs SPY $40,171 · QYLD from 2014: $22,705 vs QQQ $63,986; every dividend and income fund still ended with less money than its index · R-088 (rule: month-end close above its 10-month average, else bills; 21 themes, each vs QQQ held from the same day): beat QQQ in 0 of 21, median -9.0%/yr; boom-entry version 0 of 21, median -13.4%/yr; worst fall shallower than holding the theme in 21 of 21 · docs/research/R-087, R-088
B-RANKmeasured

Chips plus energy, held plainly beside the engine

On twenty-four years that had nothing to do with choosing it, every pair made the worst fall deeper, and none earned enough extra to pay for it.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The income funds on our watchlist are a chip fund and an energy fund with options written on top. The options only take away: 26 of 27 income funds trailed what they are built on (grave 146). So the fair question was whether the plain pair underneath works. A chip fund with an oil fund, with pipelines, with oil producers, with utilities: five pairs, each 10% of the engine, run from the year each fund existed to the end of 2024, so the chip boom that suggested the idea could not vote. Every pair earned a little more than parking the same 10% in Treasury bills, between 0.3% and 1.3% a year, and every pair paid for it with a deeper worst fall, 1.4% to 3.4% of the account. The trade was never worth it. In 2008 every one of them lost more. From the day SMH became the fund it is now, every SMH pair did worse than bills outright.

WORST FALL WITH THE PAIR, MINUS WITH BILLS-1.6%SMH/XLE+1.2%/yr-2.2%SOXX/XLE+0.6%/yr-1.4%SMH/XLU+1.3%/yr-2.7%SMH/XOP+0.7%/yr-3.4%SMH/AMLP+0.3%/yrbottom row: extra return a year for that extra fall · 10% of ULTRA, $25,000, each pair from its first year to 2024
read the autopsy

Registered before the run (R-086, commit 23b82616). The watchlist came from 336 combinations tried on one year, September 2025 to September 2026. CHPY has no other history, so only a forward test can judge it; that test is running. Its plain versions have fifteen to twenty-six years, and this used them. The shape is the watchlist’s own: 5% and 5%, carved from ULTRA’s stock side, held all the time, at $25,000. To pass, a pair had to beat the same engine with that 10% in bills and with it in gold, by a margin fixed in advance, with no deeper worst fall than the gold version, in both halves of the window, and with 2008, 2020 and 2022 each left out in turn. None came close on the first rule. The two nearest, chips with utilities and chips with oil majors, added about a point and a quarter a year, and matched it with extra swing. Switching the pair in and out on its own trend, the way every earlier sleeve was tried, did worse. Using SOXX in place of SMH before 2011 did not change a verdict. At $5,000 the pair cannot even be built: one share of SMH costs more than the whole 10% sleeve. It joins the chips (16, 121), energy (154) and the sector book (156). What is left of the income idea is one forward question, whether the option wrapper beats its own plain twin, and history says it will not.

R-086 (R-012 harness, cash model v2, whole shares, settled cash; ULTRA 90% + pair 10%, fixed, monthly rebalance; $25,000; Tiingo adjusted closes, frozen, sha256) · SMH/XLE from 2001: 16.2%/yr, worst fall -25.5% vs bills 15.0%/yr, -23.9% · SOXX/XLE from 2002: 17.1%/yr, worst fall -25.1% vs bills 16.6%/yr, -22.9% · SMH/XLU from 2001: 16.3%/yr, worst fall -25.2% vs bills 15.0%/yr, -23.9% · SMH/XOP from 2007: 19.0%/yr, worst fall -26.3% vs bills 18.3%/yr, -23.6% · SMH/AMLP from 2011: 20.2%/yr, worst fall -26.2% vs bills 20.0%/yr, -22.8% · 0 of 5 pass · docs/research/R-086
B-RANKmeasured

Gold carried on borrowed money instead of bought with the book’s own capital

Measured at the same worst fall it only helps if you already want a lot of risk; where most accounts would actually run, the interest costs more than the gold adds.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The institutions call it return stacking: keep the whole stock position and carry gold on top with borrowed money, instead of selling some stock to buy the gold. It is the one funding idea every dead sleeve before it left untested, because each of those had to be paid for out of the engine’s own capital. Carried on top, gold does look better — a point to three and a half points a year better. But that is leverage talking. Line the two up at the same worst fall and the advantage shrinks to about two thirds of a point where the book is run hottest, to nothing in the middle, and it turns negative — about four tenths of a point against — at the calmer settings. The reason is plain once you see it: selling stock to buy gold costs you return but also takes risk off the table, while borrowing costs interest whether the market rises or falls. At a high risk target the stock you gave up was worth more than the interest; at a low one the interest is the whole bill and nothing offsets it.

BORROWED GOLD MINUS BOUGHT GOLD, AT THE SAME WORST FALL+0.65DEEP END+0.08MIDDLE-0.42SHALLOW ENDmean of 24 cells · positive in 18, 16 and 5 of them · three dials, two sizes, two windows
read the autopsy

Registered twice, and the first registration was the mistake. R-070 asked a borrowed-money arm to beat a capital-funded one with no deeper worst fall. Borrowing adds exposure without removing any, so it was deeper in all twenty-four cells and no result could have passed; it also compared the two on raw return, where leverage wins by definition. That gate was written before the run and its verdict stands, but a gate nothing can pass is not a test, so the question was registered again properly as R-071 before it was run again. The second time every arm was laddered across eight dial settings so the comparison is one whole risk-return curve against another, and the places they are compared are the quarter, half and three-quarter points of the range where the two curves overlap — fixed by the data, not chosen after the fact. The overlay itself is gold notional equal to a tenth or a fifth of the account, switched on and off by the same monthly trend rule the shipped gold seat uses, financed at the three-month Treasury rate plus a third of a point, carried in its own ledger so it counts toward the account’s value and its sizing but never as cash the engine can spend on shares. Borrowed gold beats bought gold at the deep end of the range in eighteen of twenty-four cells, by two thirds of a point a year on average; in the middle it is a wash, sixteen of twenty-four and eight hundredths of a point; at the shallow end it loses in nineteen of twenty-four, by four tenths. Doubling the borrowing cost to three quarters of a point does not rescue any cell. One group of cells is honestly inconsistent and is recorded as noise rather than averaged in quietly: on ULTRA since 2009 the middle point is negative in all four cells while the shallow point is positive in all four, the opposite of every other dial, and ULTRA’s worst fall is already known to move about four points on a two per cent change in position size. Neither study ran a bitcoin version. The gold arms failed, and reaching for a second asset after a failure is how a result gets manufactured. What the record now says: the funding trick is real and it is not free, and for this book the interest is the larger bill everywhere except the hottest settings. The gold seat stays bought with capital.

R-070 and R-071 (corrected cash model, 98% sizing reserve, whole shares, $1 a leg, next-close fills; STEADY / SELECT / ULTRA, $25,000 and $250,000, 1999-03-10 and 2009-01-01 to 2026-09-01; overlay 10% and 20% of account value, gated on gold’s own monthly rule, financed at DGS3MO + 0.35% and again at + 0.75%) · stacked minus carved at matched worst fall, 24 cells: deep end +0.65 pp/yr mean, positive in 18; middle +0.08, positive in 16; shallow end −0.42, positive in 5 · ranges −0.49 to +1.65, −1.37 to +1.15, −1.09 to +1.43 · on raw return (the wrong comparison) stacking wins every cell by +0.48 to +3.50 · R-070 gate unreachable by construction (condition 2 required a levered arm to be no deeper than an unlevered one); R-071 gate +0.25 pp/yr at all three reference falls in all four window-by-size cells, and again at the doubled spread · no $5,000 arm: financed notional needs futures or margin, both gated well above it (grave 141) · docs/research/R-070, R-071
B-RANKmeasured

The levered leg as a three-times fund, two-thirds the size

Its edge was the freed cash; priced at the real bill rate and sized as the executor sizes, half a point a year was left, bought with a point of extra fall and a fund one bad day can end.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Two-thirds of the levered leg in a three-times fund, the freed third in cash: the same exposure with one fund’s fees on two borrowed dollars instead of two funds’ fees on one. On the first pass it looked like a real find: +1.47 points a year on ULTRA over the full record with a shallower fall, +0.84 on the real funds since 2010, and it survived an attack of eras, years, taxes, crash sessions and costs. Then the checks that model the account as it is actually run took it apart, one accounting detail at a time: the executor’s 2% sizing reserve, real funds at real sizes, and finally interest at the historical bill rate instead of a flat 3%. At $25,000 the edge ended at +0.47 of a point with the worst fall 0.88 points deeper; at $5,000, +0.13. Below its own pass mark on both counts.

THE THREE-TIMES FUND ON ULTRA: EXTRA RETURN (GREY) AND EXTRA FALL (RED), EACH TEST ROUND+1.47+1.64MODEL 1999-+0.84+1.11REAL 2010-+0.70-0.90+RESERVE+0.47-0.88+REAL CASHgrey extra return, points a year · red extra worst fall, points (negative = shallower) · ULTRA, $25,000
read the autopsy

Registered as R-066 with the pass mark written first: at least a quarter point a year on both the modelled full window and the real funds, at both sizes, with no deeper worst fall (half a point of tolerance). The first run passed on ULTRA (+1.47 / +1.52 modelled, +0.84 / +1.05 real, falls -28.3% to -26.7%) and failed on SELECT and STEADY; R-068 confirmed the S&P leg gains nothing worth the fall either. R-067 attacked it: positive at equal fall in all three eras (+0.3, +2.1, +2.8), 18 of 27 years, +1.4 after tax, one to four points deeper in a crash session, +0.9 at $3,000. The Lab replayed it independently, 56 simulations, and reproduced every number. Then the Lab asked for the account as it is run, not as the harness idealises it. With the executor’s 98% sizing reserve the real-fund edge at $25,000 fell to +0.70 with the fall 0.9 points deeper, and at $5,000 to +0.07: whole shares and a 2% smaller slice move both arms enough that the difference is noise at small sizes. The Lab then found that the simulator pays parked cash a flat 3% whatever it is told, and rebuilt that line to accrue the historical three-month bill rate on settled and parked balances (eight accounting tests, sixteen exact legacy comparisons). Under that model the edge at $25,000 is +0.47 points (23.3% against 23.8%) with the fall 0.88 points deeper (-26.5% to -27.4%), and at $5,000 +0.13 with 0.78 deeper. The return clears the quarter point at $25,000 and not at $5,000; the fall clears the tolerance at neither. What the record shows: the three-times fund’s advantage is mostly the freed cash, and the freed cash is worth what cash earns; price it at the historical rate, size it as the executor sizes it, and hold it in whole shares at the sizes people actually run, and about a quarter to half a point a year is left, bought with a point of extra fall and a fund that a single 33% session would end. Two things worth more than the idea came out of it: the published record’s cash accrual is a flat 3% that the site now needs to reconcile in a controlled update, and ULTRA’s worst fall moves by up to four points on a 2% sizing change, a path sensitivity the record should state. The levered leg stays QLD on every dial.

R-066 (modelled 3x = 3r minus two borrowed dollars at fed funds + 0.35% minus 0.84% expense; real QLD vs TQQQ 2010-04..2026-08; $25k and $250k): ULTRA modelled full +1.47 / +1.52 pp, falls -28.3% → -26.7%; real +0.84 / +1.05; SELECT +0.03 (fall deeper), STEADY +0.27 (fall deeper) · R-067: equal-fall edge by era +0.31 / +2.14 / +2.75; 18 of 27 years; after tax +1.38; sizes $3k +0.92, $5k +1.25, $25k +1.47 · R-069 with the 98% reserve and the live cash rules, real funds: $25k +0.70 (fall -26.5% → -27.4%), $250k +0.95 · the Lab’s cash model v2 (historical three-month bill accrual, DGS3MO; outputs/r066-audit/cash-v2/REPORT.md): $5k QLD 23.13% vs TQQQ 23.26%, +0.134 pp, falls -26.57% → -27.35%; $25k 23.29% vs 23.75%, +0.466 pp, falls -26.49% → -27.37% · gate: +0.25 pp both windows both sizes, fall no deeper than 0.5 pt · docs/research/R-066..R-069
B-RANKmeasured

Reading the bitcoin seat every day instead of once a month

A quarter-point lead on the years it was picked on became nothing on the years it was judged on; it was a fit to the 2018 crash.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The bitcoin seat looks at bitcoin once a month against its 200-day average. A grid the day before said a 100-day average read every session did better on every window we had, so we did it properly: picked the setting on 2015 to 2019, judged it on 2020 to 2026, and charged every trade it makes. On the years it was picked on it led by +0.25 Sharpe on ULTRA. On the years it was judged on it led by -0.028 on ULTRA and +0.030 on SELECT, with two to three times the trades. An edge that shrinks from a quarter of a point to nothing when it meets years it was not fitted to is not an edge. The monthly read stays.

A FASTER BITCOIN READ: SHARPE OVER THE SHIPPED RULE, PICKED ON 2015-19 (GREY), JUDGED ON 2020-26 (RED)+0.25-0.03ULTRA+0.23+0.03SELECT+0.26+0.02STEADYgrey the years it was picked on · red the years it was judged on · $25,000, real trading costs charged
read the autopsy

R-063 ran the gold seat’s plateau check on the bitcoin seat and found the short-and-fast corner of the grid beating the shipped rule on both windows at both sizes: the 100-day average read every session, +0.06 to +0.15 Sharpe beside the engine. Three objections were written down before anything moved: it was the corner of the grid, it was picked and judged on the same data, and the seat stream charged no trading cost. R-065 answered all three, registered before the run: lengths 50 to 200 days, three cadences, the seat traded by the harness itself (whole units, a $1 commission per fill, next-close fills, T+1 settled cash, the engine leg untouched), the candidate chosen on 2015 to 2019 by its mean lead across the three dials and judged on 2020 to 2026 at $25,000 and $250,000, pass mark +0.03 at both sizes on ULTRA and SELECT with the worst fall no deeper by more than two points. Training chose the same cell R-063 had, 100/D, by a wide margin: +0.252 on ULTRA, +0.235 on SELECT, +0.259 on STEADY; every daily-read cell led by 0.15 or more, every weekly one by about 0.10, so the training years were unanimous. The test years were not: ULTRA -0.028 at $25,000 and +0.004 at $250,000, SELECT +0.030 and +0.012, STEADY +0.017 and +0.007; the book earned 28.8% a year against 30.6% on ULTRA with a worst fall 2.1 points shallower, and the seat traded 12 times a year against 5. What the training years had that the test years did not: 2018, when bitcoin fell 80% over eleven months and a faster read left two months earlier than a monthly one. 2022 was the same kind of fall and the faster read gained a little there too, but 2020 to 2026 also held March 2020 and a run of whipsaws in 2021 and 2024 where a daily read stepped out and back in at a cost, and the two effects cancel. The shallower worst fall is real and small, the extra return is not, and the extra trades are certain. The monthly rule stays, the bitcoin gate stays on its clock, and the corner cell of R-063 is recorded as what it was: a fit to one crash.

our full test on the R-012 harness with the seat traded by the simulator (scripts/research/r065/recompute_btc.py), bitcoin as BTC-USD on equity sessions, 20% seat all in or out, engine leg as shipped; training 2015-06-01 to 2019-12-31 at $25,000, test 2020-01-02 to 2026-08-31 at $25,000 and $250,000 · training mean over shipped: 100/D +0.249, 75/D +0.212, 150/D +0.183, 200/D +0.183, 50/D +0.148, 50/W +0.129 · test, candidate 100/D over shipped (Sharpe; fall deeper pts; seat trades/yr shipped→candidate): ULTRA $25k -0.028 / -2.1 / 4.67→11.9; ULTRA $250k +0.004 / -3.1 / 4.67→11.75; SELECT $25k +0.030 / -2.8 / 7.98→13.86; SELECT $250k +0.012 / -2.8 / 8.13→13.86; STEADY $25k +0.017 / -2.2 / 6.93→13.41; STEADY $250k +0.007 / -1.9 / 7.38→13.41 · full window $25k shipped / candidate: ULTRA 32.0% / 34.2% (Sharpe 1.35 / 1.46), SELECT 29.3% / 30.7%, STEADY 27.5% / 29.1% · docs/research/R-063-results.md, R-065-results.md
B-RANKmeasured

A permanent long-volatility seat

Insurance that pays once a decade and bleeds the other nine years; since 2018 it looks brilliant because 2020 is inside the window and 2011 to 2017 is not.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Crash insurance a retail account can buy without options: a small seat in a fund that holds volatility futures. On the full window since 2011 every construction lost to the gold sleeve on every dial. Start the clock in 2018 and the mid-term fund as the whole seat beats gold by +0.12 Sharpe on SELECT and +0.18 on STEADY. The difference is one year: the fund made +76% in 2020 and lost -92% from 2011 to 2017, when there was nothing to insure. A window that starts in 2018 keeps the payoff and forgets the premium. The full window is the verdict.

THE MID-TERM VOLATILITY FUND AS THE WHOLE SEAT, MINUS THE GOLD SLEEVE: FULL WINDOW vs SINCE 2018-0.08+0.02ULTRA-0.08+0.12SELECT-0.06+0.18STEADYgrey 2011 to 2026 (the verdict) · red 2018 to 2026 (2020 inside, 2011 to 2017 outside) · $25,000
read the autopsy

Mel asked whether the VIX playbook had been tested. The option hedge (graves 84, 118), every VIX signal (43, 85, 106, 129, 130, 131) and the short side of the volatility funds (102) already had stones; the long side as a small permanent seat did not. Registered as R-061: VIXY (short-term futures), VIXM (mid-term) and VXX, from 2011 and since 2018, four ways in the seat: the fund as the whole seat, a quarter of the seat (5% of the book, the size a hedge would be), that 5% only while the engine is in, and that 5% only while the fund is above its own 200-day; against cash, gold and gold-while-out; pass = +0.03 Sharpe at both sizes with no deeper fall. The expectation was written first: the short-term fund’s roll cost makes it a certain loser, the mid-term fund is the one cell with a chance. Alone: VIXY -48.6% a year and a -100% worst fall, VIXM -17.6% and -96%; correlation with the Nasdaq fund -0.70; inside the 2022-23 stretch the engine sat out, the funds lost -49% and -23% while gold made +7%: a slow bear is not a volatility event. Full window, best construction per dial against gold at $25,000, VIXM: ULTRA -0.004, SELECT -0.040, STEADY -0.011; VIXY worse on every line. Since 2018 the mid-term fund as the whole seat passes SELECT and STEADY at both sizes (book 15.2% a year, -16% worst fall on SELECT) and the 5% short-term seat held while the engine is in passes STEADY. Why that is the trap, not the find: the validation window must contain the failure mode, and a long-volatility seat’s failure mode is a long calm bull market with nothing to insure. 2011 to 2017 is exactly that, and the mid-term fund lost -92% across it, about -32% a year of the seat, before 2018 gave it February, December and then the 2020 crash (+76%). The since-2018 pass is one payoff with its premium cut off. The engine’s exit insures the same event for the cost of being a month late, and gold pays in the slow bears where volatility funds bleed.

our full test on the R-012 harness, v2 sleeve machinery, 2011-04-04 to 2026-09-11 and since 2018-01-02, $25,000 and $250,000, shipped rules · alone: VIXY -48.6% / -100% / Sharpe -0.62, VIXM -17.6% / -96% / -0.44, VXX (2018-04) -45.4% / -100%; VIXM 2011-04..2017-12 -92.1%, 2018-01..2026-09 -35.9%, 2020 +76.3%; VIXY 2011-17 -99.6%, 2020 +15.0% · seat over gold, full window $25k / $250k, VIXM: vol-full ULTRA -0.075/-0.106, SELECT -0.079/-0.076, STEADY -0.065/-0.041; vol-5pct ULTRA -0.035/-0.048, SELECT -0.041/-0.036, STEADY -0.015/-0.023; vol-5pct-while-in ULTRA -0.004/-0.052, SELECT -0.040/-0.037, STEADY -0.011/-0.016; vol-5pct-gated ULTRA -0.038/-0.085, SELECT -0.047/-0.055, STEADY -0.027/-0.014 · since 2018, VIXM vol-full over gold: SELECT +0.119/+0.084, STEADY +0.180/+0.149, ULTRA +0.024/+0.066; VIXY vol-5pct-while-in STEADY +0.041/+0.040 · verdict on the full window, which holds the 2011 to 2017 failure mode · docs/research/R-061-results.md
B-RANKmeasured

Emerging markets switched on by a falling dollar

A falling dollar is a risk-on regime, not an emerging-markets regime; the Nasdaq the engine already owns rises more in the same months.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The case for emerging markets is a dollar case: they rise when the dollar falls and sink when it rises. That part is true. From 2004 the emerging-markets fund earned 11.1% a year while the dollar sat below its 200-day average and 4.5% while above it. So we made the dollar the seat’s switch. It still lost to the gold sleeve on every dial, both sizes, five ways. The contrast row is the autopsy: the S&P 500 shows the same split (12.8% against 8.7%) and the Nasdaq a bigger one (19.3% against 10.2%). A falling dollar is not an emerging-markets regime. It is a risk-on regime, and the engine already owns the best risk-on asset there is.

RETURN PER YEAR WHEN THE DOLLAR IS FALLING (GREY) AND RISING (RED), 2004-2026+0.11+0.05EMERGING+0.13+0.09S&P 500+0.19+0.10NASDAQgrey dollar below its 200-day · red above · return per year, 0.20 = 20%
read the autopsy

Mel asked for reasoning rather than another fund, and the reasoning gave a mechanism: emerging-market earnings, funding and flows are priced in dollars, so their good decades (2003 to 2007, and 2017) sit inside dollar downtrends and their bad ones (2011 to 2015, 2018, and 2021 to 2022) inside dollar uptrends. Grave 150 had already killed emerging markets held outright or on their own trend; this study switched the seat on the dollar’s trend instead: the dollar index’s last close against its 200-day average, read at the month turn, hold EM while the dollar is below, cash while above. Registered first, with a mechanism check written in: if the split by regime is not there, the seat result is moot. The split is there. EEM 11.1% a year and Sharpe 0.58 below the average against 4.5% and 0.30 above; VWO the same. But so is everything else: the S&P 500 12.8% against 8.7%, the Nasdaq-100 19.3% against 10.2%, at higher Sharpe in both regimes. The dollar falls when the world is buying risk, and the engine’s own root is the highest-return risk asset in the set, so the switch adds a slower, lower-return copy of a bet the book already holds. The seat streams alone: EM always 7.8% a year with a -66% worst fall; dollar-gated 5.8% and -42%; both gates 5.4% and -31%; the gate cuts the fall and the return together. Beside the engine the best construction sits -0.099 (ULTRA), -0.093 (SELECT), -0.086 (STEADY) Sharpe below the gold sleeve at $25,000 on the full window, and no better at $250,000 or since 2009; the only cells above cash are the while-out ones since 2009, which is the gold-while-out seat with a worse asset. Uranium and the nuclear theme in the post that prompted this were not retested: URA is inside grave 150, and a supply-timing narrative with an 80% drawdown in its own history (2011) is not a rule the engine can read.

our full test on the R-012 harness, v2 sleeve machinery, EEM 2004-02-08 to 2026-09-11 and since 2009-01-02, VWO from 2006-01-04, $25,000 and $250,000, shipped rules; dollar = ICE dollar index daily closes; gate read at the month turn · mechanism: EEM below/above 11.1% / 4.5% (Sharpe 0.58 / 0.30, 3000 / 2709 sessions); VWO 10.4% / 1.3%; SPY 12.8% / 8.7% (0.88 / 0.49); QQQ 19.3% / 10.2% (1.03 / 0.52) · seat alone (EEM): em-always 7.8% / -66% / 0.41; em-own-gate 6.7% / -48% / 0.45; em-dollar-gate 5.8% / -42% / 0.42; em-dollar-and-own 5.4% / -31% / 0.42; em-dollar-gate-while-out 1.1% / -24% / 0.25 · best over gold per dial, full window, $25k / $250k: ULTRA -0.099 / -0.092, SELECT -0.093 / -0.104, STEADY -0.086 / -0.093; since 2009 ULTRA -0.072 / -0.061 · docs/research/R-060-results.md
B-RANKmeasured

Miners or two-times gold as the gold seat’s holding

Miners are an equity with a gold accent, and a doubled holding under a capped, vol-sized seat is halved back with the fund’s drag left over.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The Nasdaq leg reads the index and holds the two-times fund. Could the gold seat do the same, read gold and hold something that moves more than gold: miners, silver miners, junior miners, or the two-times gold fund? Under the shipped gate none beat the seat as shipped on any dial at either size. Miners are an equity with a gold accent: -20.5% in the 2008 stretch the engine sat out while gold made -1.6%, and worst falls of -81% to -89% on their own. The two-times fund came closest, within a hair on two dials, and paid for it in every calm year.

MINERS OR TWO-TIMES GOLD IN THE GOLD SEAT, BEST READ, MINUS THE SEAT AS SHIPPED-0.07-0.09GDX-0.02-0.03SIL-0.03-0.07GDXJ-0.03+0.04UGLgrey $25,000 · red $250,000 · ULTRA · best of read-on-gold / read-on-itself
read the autopsy

Registered as R-059 part B before the run: the gold seat’s gate and cap as shipped (last close against the 200-day average at the month turn, vol-sized, capped at 1.0) with the holding swapped for GDX (from 2006-05-22), SIL (2010-04-20), GDXJ (2009-11-11) and UGL (2008-12-03), read on gold’s signal and read on the vehicle’s own closes, v2 sleeve machinery, three dials, $25,000 and $250,000, each from its listing, against the shipped gold seat over the same window; pass = +0.03 Sharpe at both sizes with no deeper fall. Alone: GDX 5.6% a year with a -81% worst fall, SIL 2.9% / -83%, GDXJ 1.8% / -89%, UGL 10.4% / -76%; correlation with gold +0.76 to +1.00. In the seat the best read per vehicle sits below the shipped seat on ULTRA by -0.074 (GDX), -0.016 (SIL), -0.035 (GDXJ) and -0.028 (UGL) Sharpe at $25,000, with SELECT and STEADY worse for the miners. Why the analogy with QLD fails: the two-times Nasdaq fund is the index with financing; miners are companies with costs, debt and an equity beta, so they fall with stocks in the stretches the seat exists to cover, and the gate, which reads gold, cannot see that. The two-times gold fund is the honest analogue and it lost anyway: the seat is capped at 1.0 and vol-sized, so doubling the holding’s volatility halves its size and adds the fund’s reset drag and expense for nothing. The gold seat stays SGOL.

our full test on the R-012 harness, v2 sleeve machinery, each vehicle from listing (+300 days) to 2026-09-11, $25,000 and $250,000, shipped gate and cap · alone: GDX 5.6% / -81% / Sharpe 0.34 / corr gold +0.77, 2020 +23.7%, 2022 -9.0% · SIL 2.9% / -83% / Sharpe 0.27 / corr gold +0.73, 2020 +40.3%, 2022 -22.8% · GDXJ 1.8% / -89% / Sharpe 0.27 / corr gold +0.76, 2020 +30.4%, 2022 -14.5% · UGL 10.4% / -76% / Sharpe 0.46 / corr gold +1.00, 2020 +39.0%, 2022 -7.6% · GDX 2008 -26.1%; engine-out 2008-09 GDX -20.5% vs gold -1.6% · Sharpe over the shipped seat, best read, $25k / $250k: GDX ULTRA -0.074 / -0.091, SELECT -0.105 / -0.100, STEADY -0.123 / -0.105; SIL ULTRA -0.016 / -0.032, SELECT -0.071 / -0.074, STEADY -0.080 / -0.097; GDXJ ULTRA -0.035 / -0.072, SELECT -0.087 / -0.075, STEADY -0.120 / -0.119; UGL ULTRA -0.028 / +0.041, SELECT +0.035 / +0.013, STEADY +0.005 / +0.012 · docs/research/R-059-results.md
B-RANKmeasured

The two-times leg financed by a box spread

Cheaper financing is real but small, only ULTRA borrows enough to notice, and the futures engine already gets it with less hassle.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Grave 141 killed borrowing from the broker to hold the plain fund instead of the two-times fund: the broker charges too much. A box spread, an options trade that borrows at close to the Treasury rate, is cheaper than the fund’s own financing by about half a point per borrowed dollar. Under the shipped rules that is worth +0.51 points a year on ULTRA at $25,000 and +0.03 on SELECT, which borrows little. The pass mark asked for a quarter point on both dials; it cleared one. And the futures version of the engine already gets that cheaper financing at $25,000 and up without options permissions, boxes to roll, or $10,000 borrowing steps.

THE TWO-TIMES LEG ON BOX FINANCING MINUS THE FUND, RETURN PER YEAR+0.03+0.02SELECT+0.51+0.27ULTRAgrey $25,000 · red $250,000 · points a year, box at 0.35% over fed funds
read the autopsy

The two-times fund carries its own financing: fed funds plus about 0.35% of swap spread plus a 0.95% expense ratio, 1.3% over fed funds per borrowed dollar. R-033 (grave 141) showed a margin loan at the broker’s 1.5% spread costs more, not less. A box spread on cash-settled index options is a synthetic loan priced by arbitrage; the published record puts its rate about 0.35% over the matching Treasury rate on average. Holding the plain fund on box financing therefore costs fed funds plus 0.35% plus twice the plain fund’s 0.20% expense: 0.75% over fed funds, a saving of 55 basis points per borrowed dollar. We registered the test as R-033’s harness with the cost model swapped, three spreads (0.20%, 0.35%, 0.80%), SELECT and ULTRA at $25,000 and $250,000, 1999 to 2026, and the gate: at least a quarter point a year on both dials at both sizes at the baseline spread, surviving the 0.80% stress at one size. Result at 0.35%: SELECT +0.03 / +0.02, ULTRA +0.51 / +0.27 points; at 0.80% ULTRA keeps +0.27 at $25,000 and +0.05 at $250,000; worst falls unchanged. The mechanism is exactly what the cost gap predicts: the saving scales with how much is borrowed, ULTRA (cap 2.0) borrows about twice what SELECT does, and fed funds averaged 2.14% over the window so the whole financing line was small. Real but small, and the registered gate wanted both dials. Not modelled, as in R-033: the path difference between a daily-reset fund and a loan (the band keeps the two close), one-box granularity ($10,000 of borrowing per SPX box, a fifth of a $25,000 ULTRA account’s exposure), quarterly rolls, and the options permission and margin account a box needs. The practical verdict: the futures engine from $25,000 already borrows at futures rates, about 0.2% over fed funds, with one contract and no options; a box spread is a harder way to the same place.

our full test on the R-012 harness, R-033 cost model, 1999-03-10 to 2026-08-31, shipped rules, SELECT and ULTRA, $25,000 and $250,000 · fund cost fed funds + 130 bp per borrowed dollar; box cost fed funds + spread + 40 bp · return/yr difference, box minus fund, SELECT $25k / $250k, ULTRA $25k / $250k: 0.20% spread +0.04 / +0.04, +0.59 / +0.34; 0.35% +0.03 / +0.02, +0.51 / +0.27; 0.80% -0.02 / +0.02, +0.27 / +0.05 · control SELECT 15.54% / -20.3%, ULTRA 17.24% / -28.2% at $25k · no cell with a worse fall · average fed funds 2.14% · docs/research/R-059-results.md
B-RANKmeasured

An ETF breakout book in the seat

A faster trend rule on the same funds is the engine’s own signal with more noise; it never beat the gold sleeve.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The other old shadow book: buy any of twenty-eight big funds that closes at a new 55-day high, sell at a 20-day low, a stop, or after thirty days, five at a time. Rebuilt from its rule and run as a fifth of the book beside every dial. Alone it earned 7.4% a year with a -24% worst fall and won 41.7% of its trades; in the seat it lost to the gold sleeve on every dial, every way it was held, at both sizes. A trend rule on funds cannot help a trend engine on the same funds: it is out when the engine is out and in when it is in.

THE BREAKOUT BOOK IN THE SEAT, BEST OF FOUR WAYS, MINUS THE GOLD SLEEVE-0.07-0.08ULTRA-0.08-0.08SELECT-0.07-0.08STEADYgrey $25,000 · red $250,000 · best of always / gated / while out / both
read the autopsy

The Donchian quant book (“DON55/20”) ran as a report-only shadow on our box and had already halted itself by its own floor (51 closed trades, -0.16R). Its rule, rebuilt from the published definition: signal when the close is at or above the prior 55-day high, fill at the next open, stop at two ATR below entry, exit on the first close at or below the prior 20-day low once twenty bars have been held, at the stop, or at the close of the thirtieth bar; five positions at most, strongest breakout first; 1.5 bp a side. Registered with the allocator in R-058 before the run. As a five-slot seat from 2005: 7.4% a year, worst fall -24%, Sharpe 0.69, 50 trades a year, 41.7% wins, correlation +0.35 with the Nasdaq fund; -2.8% in 2008, -11.3% in 2022. Inside the engine’s out-stretches it made +7.0% in 2008-09 and lost -5.6% in 2022-23 while gold made +7.1%. Beside the engine the best construction per dial sits -0.070 (ULTRA), -0.080 (SELECT), -0.074 (STEADY) below the gold sleeve at $25,000 and no better at $250,000; it beats cash on SELECT by a hair and nowhere else. The mechanism is the one the graveyard already knows from the exit vote and the shorter averages (graves 142 to 145, R-057): a second, faster trend rule on the same market is the engine’s own signal with more noise. Where it holds something the engine does not (bonds, gold, oil, foreign funds) it is the R-053 sector book again, and those pieces did not pay in the stretch that mattered.

our full test on the R-012 harness, v2 sleeve machinery, 2005-04-01 to 2026-09-11, $25,000 and $250,000, shipped rules; 28-fund universe, yfinance adjusted bars from 2004 frozen under scripts/research/r058/inputs · alone: 7.4% / -24% / Sharpe 0.69, 49.6 trades/yr, 41.7% wins, corr QQQ +0.35; 2008 -2.8%, 2020 +26.0%, 2022 -11.3% · engine-out 2008-09 +7.0% (gold -1.6%, cash +1.4%), 2022-23 -5.6% (gold +7.1%) · best over gold per dial $25k / $250k: ULTRA -0.070 / -0.080, SELECT -0.080 / -0.076, STEADY -0.074 / -0.077 · the rule is a re-implementation of the shadow’s stated parameters, not its code · docs/research/R-058-results.md
B-RANKmeasured

An ETF dip-buying book in the seat

It buys the dips the engine already owns, needs same-day cash the engine keeps for its adds, and trades its return away in commissions.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

An old shadow book of ours: every close it ranks twenty-two big funds by how hard they fell over the last few days, buys the two most oversold at the next open, and sells on the first bounce, a stop, or the fifth day. Rebuilt from its rule and run as a fifth of the book beside every dial. On its own it looks fine: 9.2% a year, worst fall -19%, up +18.9% in 2008. In the seat it never beat the gold sleeve, on any dial, held any of four ways, at either size. And it trades 222 times a year, which at a one-dollar commission on a $5,000 seat costs about 9% a year before it has earned anything. Retested for Robinhood (R-146, 1 October 2026): with the engine’s own commissions set to zero it still trailed the gold slice on every dial and at both account sizes. The fees were never what stopped it.

THE DIP-BUYING BOOK IN THE SEAT, BEST OF FOUR WAYS, MINUS THE GOLD SLEEVE-0.04-0.05ULTRA-0.04-0.03SELECT-0.04-0.04STEADYgrey $25,000 · red $250,000 · best of always / gated / while out / both
read the autopsy

The ETF cross-sectional mean-reversion allocator ran as a report-only shadow on our box from July 2026 (29 forward trades, +0.24R, 72% wins) after a July study of its own found it failed four of five validation gates and lost money at a $5,600 account on the broker’s fixed commission. Mel asked whether anything in the retired shadows could be a seat, then said test all. We wrote the rule and the pass mark down first. The rule was rebuilt from its published definition (not the retired code): z-scores across the universe of the 3- and 5-day returns and the distance to the 5- and 10-day averages in volatility units, plus a two-day RSI term; top two above a threshold, one new position a day, five at most, next-open fills, a 1.5 ATR stop, exit on the first close above the 5-day average or the fifth day; 1.5 bp a side. The rebuild put the shadow’s own pick in its top three on 27 of 29 forward days (the exact pick on 17, the rest rank swaps at near-equal scores and the shadow’s empty July book). As a five-slot seat with idle slots in cash it earned 9.2% a year from 2005 with a -19% worst fall and a Sharpe of 0.96, +18.9% in 2008 and -5.7% in 2022, correlation +0.57 with the Nasdaq fund. Inside the engine’s out-stretches: +25.6% in 2008-09 against gold -1.6%, +4.4% in 2022-23 against gold +7.1%. Beside the engine it beats cash on SELECT (+0.066 Sharpe held always) and loses to the gold sleeve everywhere: the best construction per dial sits -0.040 (ULTRA), -0.039 (SELECT), -0.037 (STEADY) below gold at $25,000, no better at $250,000. Two reasons it cannot be a seat even where it edges cash: it is long the same dips the engine already owns through the two-times fund, so it adds exposure to the engine’s own bad days rather than the opposite; and it needs same-day cash on a settled-cash account that the engine keeps liquid for its own adds. The commission arithmetic is the third: 222 round trips a year on a $5,000 seat is about $450 at the fixed minimum, nine points of a nine-point return.

our full test on the R-012 harness, v2 sleeve machinery, 2005-04-01 to 2026-09-11, $25,000 and $250,000, shipped rules; 22-fund universe, yfinance adjusted bars from 2004 frozen under scripts/research/r058/inputs · alone: 9.2% / -19% / Sharpe 0.96, 222.3 trades/yr, 65.1% wins, corr QQQ +0.57; 2008 +18.9%, 2020 +16.6%, 2022 -5.7% · Sharpe over cash / gold / gold-while-out, $25k, held always: ULTRA +0.022 / -0.040 / +0.018; SELECT +0.066 / -0.039 / +0.058; STEADY -0.037 / -0.037 / +0.040 · best over gold per dial $25k / $250k: ULTRA -0.040 / -0.049, SELECT -0.039 / -0.029, STEADY -0.037 / -0.037 · parity with the shadow’s forward ledger: top-3 27/29, exact 17/29 · docs/research/R-058-results.md
B-RANKmeasured

Fractional shares to open the engine below $3,000

Whole shares cost nothing; the $1 commission is what taxes a small account.
small account or this brokerTested as accounts under $3,000.

We removed the share rounding from the simulator and ran every dial from $1,000 to $25,000 over the full cycle. Fractional fills moved the return by at most 0.53 of a point a year in either direction, and whole shares came out slightly ahead more often than not. What separates a $1,000 account from a $25,000 one is the $1 commission: 2.4 points a year on STEADY at $1,000, 0.4 at $5,000, 0.07 at $25,000. A commission-free account closes that gap; fractional shares do not.

BELOW $25,000: WHAT SHARE ROUNDING COSTS vs WHAT THE $1 COMMISSION COSTS, STEADY-0.53-2.42$1k-0.04-1.10$2k-0.06-0.72$3k-0.03-0.41$5k-0.07-0.17$10k-0.08-0.07$25kgrey what fractional fills add · red what the $1 commission costs · STEADY, points a year
read the autopsy

The pitch was a product rung, not a return: if rounding to whole shares is what makes a $2,000 account drift from the published record, fractional orders, which the broker fills on liquid funds, would open the engine below its floor for free. We registered the two conditions a rung had to meet before running: rounding must cost at least 0.30 of a point a year, and the fractional account must then land within 0.30 of the $25,000 record. Neither held anywhere. Rounding cost between -0.53 and +0.38 of a point (a negative number means whole shares did better, because rounding down holds a little less than the target and the target is a volatility guess, not a truth). The small-account gap is the commission. STEADY places the most orders (three legs), so at $1,000 its $1 fills cost 2.4 points a year and the account earned 10.7% against 12.6% at $25,000; ULTRA, one leg, lost 0.47 at $1,000 and nothing above $5,000. With commissions at zero every size sits within a tenth of a point of the $25,000 line. So the rung question is a broker question: a commission-free account runs the engine at $1,000 as well as at $25,000 in this model, and a $1-minimum account should not run STEADY below about $5,000. The product keeps whole shares; the buyer guide gets the commission sentence instead.

our full test on the R-012 harness, frozen rules, 1999-03-10 and 2009-01-02 to 2026-08-31, whole shares vs the same simulator with only the rounding removed (scripts/research/r056/recompute_frac.py; fractional orders under $5 not sent) · STEADY full window, whole / fractional / commission-free return per year: $1,000 10.7% / 10.1% / 12.6%, $2,000 11.5% / 11.5% / 12.6%, $3,000 11.9% / 11.8% / 12.6%, $5,000 12.2% / 12.2% / 12.6%, $10,000 12.5% / 12.4% / 12.6%, $25,000 12.6% / 12.5% / 12.6% · ULTRA: $1,000 16.3% / 16.4% / 16.9%, $2,000 16.9% / 16.6% / 16.8%, $3,000 17.0% / 16.7% / 16.8%, $5,000 16.9% / 16.7% / 16.8%, $10,000 16.9% / 16.8% / 16.8%, $25,000 16.8% / 16.8% / 16.8% · SELECT: $1,000 14.2% / 14.6% / 15.4%, $2,000 14.6% / 15.0% / 15.4%, $3,000 15.1% / 15.1% / 15.4%, $5,000 15.3% / 15.2% / 15.4%, $10,000 15.2% / 15.3% / 15.4%, $25,000 15.3% / 15.3% / 15.4% · worst falls unchanged within 2 points at every cell · docs/research/R-056-results.md
B-RANKmeasured

The engine’s rules on a momentum portfolio instead of the Nasdaq

Momentum stocks fall like the Nasdaq and recover slower; the exit reads them worse.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The same exit and dial, read on the strongest tenth of the market by recent momentum instead of the Nasdaq-100. Raw, the momentum series looks better than the Nasdaq: 12.6% a year against 10.7% with a shallower worst fall. Under the engine’s rules every dial lost: ULTRA earned 12.6% a year against 16.7% with a deeper worst fall (-38% against -28%), and on the real fund since 2013 it earned 17.4% against 22.2%.

THE SAME RULES ON A MOMENTUM ROOT, SHARPE MINUS THE SHIPPED DIAL-0.19-0.14ULTRA-0.26-0.23SELECT-0.11-0.09STEADYgrey with a modelled 2x leg · red the plain fund, no leverage · $25,000, 1999-2026
read the autopsy

Cross-sectional momentum, buying what rose most over the past year, is the best-documented anomaly in the academic record and has a hundred years of data. Since 1999 the top momentum tenth returned 12.6% a year with a -56% worst fall against the Nasdaq fund’s 10.7% and -83%, and it was flat from 2000 to 2002 while the Nasdaq halved. That is the case for it as the engine’s root. We registered the test first: the frozen rules on the momentum series with a modelled two-times leg (no such fund exists), and on the plain series capped at one times exposure (what a cash account could hold in MTUM), against the shipped dials, at $25,000 and $250,000, 1999 to 2026; then the same on MTUM itself, the real fund, since its 2013 listing. Every arm lost on every dial. ULTRA: 12.6% / -38% against the shipped 16.7% / -28%; SELECT 10.7% against 15.2%; STEADY 11.0% against 12.5%. On the real fund the gap is wider: ULTRA 17.4% and 13.1% against 22.2%. Why the raw advantage vanishes under the rules: the exit and the dial are built for a series that falls hard and recovers hard. Momentum stocks fall with the market but recover slowly, because the portfolio is rebuilt from last year’s winners just as the losers rip; in 2009 the Nasdaq made +54.7% and momentum +7.1%, so the engine came back in for +34.2% on the Nasdaq and +15.7% on momentum. The volatility dial then sizes momentum’s calm stretches up and its crashes arrive from calm. The academic series is also an upper bound: monthly-formed, no costs, microcaps in; the real fund earned less on every line. Noted against ourselves: a two-times momentum fund does not exist, so the only buyable arm was the plain one, and it lost by more.

our full test on the R-012 harness, frozen rules, 1999-03-10 to 2026-07-31 (Ken French daily decile 10, value-weighted), $25,000 and $250,000; MTUM window 2013-04-18 to 2026-07-31, MTUM tracks the decile at 0.91 · roots alone: momentum 12.6% / -56% / Sharpe 0.58, QQQ 10.7% / -83% / 0.51 · ULTRA: shipped 16.7% / -28% / 0.83, mom-2x 12.6% / -38% / 0.64 (-0.191), mom-1x 11.2% / -40% / 0.69 (-0.141); MTUM window shipped 22.2%, mtum-2x 17.4% (-0.125), mtum-1x 13.1% (-0.052) · SELECT: shipped 15.2% / -20% / 0.90, mom-2x 10.7% / -32% / 0.64 (-0.262), mom-1x 10.2% / -32% / 0.67 (-0.229); MTUM window shipped 20.0%, mtum-2x 15.3% (-0.171), mtum-1x 13.9% (-0.088) · STEADY: shipped 12.5% / -18% / 0.95, mom-2x 11.0% / -18% / 0.84 (-0.113), mom-1x 10.6% / -18% / 0.86 (-0.090); MTUM window shipped 15.6%, mtum-2x 14.0% (-0.061), mtum-1x 13.2% (-0.019) · $250,000 within 0.5pp of every $25,000 line · docs/research/R-055-results.md
B-RANKmeasured

A trend-following fund in the seat, always, gated, or only while out

Trend funds had already made their 2022 before the exit turned off; beside the engine they trail gold every way they were held.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Four managed-futures funds, the oldest reaching back to 2007, held in the 20% seat four ways: always, on their own trend gate, only while the engine had stepped out, and both. Sixteen constructions on three dials at two account sizes; none beat the gold sleeve, and the single best of them, one size on one dial, cleared gold by +0.03 Sharpe while its other size did not. Inside the stretch the engine sat out in 2022 and 2023, the four funds returned -12.8% to -0.0% while gold made +7.1%.

A TREND FUND IN THE SEAT, BEST OF FOUR WAYS, MINUS THE GOLD SLEEVE, ULTRA-0.07-0.06RYMFX-0.04-0.07AQMIX-0.03-0.09DBMF-0.05+0.03KMLMgrey $25,000 · red $250,000 · best of always / gated / while out / both
read the autopsy

Managed-futures funds ride trends in many markets at once, long and short, and are sold as crisis insurance: they made money in 2008 and in 2022 while stocks fell. Both facts hold in the raw funds. The question was whether that pays beside an engine that already steps out of the market on its own trend rule. We wrote the rules and the pass mark down first, then ran RYMFX (a 2007 fund on the S&P trend index), AQMIX (AQR, 2010), DBMF (2019) and KMLM (2020) in the 20% seat on ULTRA, SELECT and STEADY at $25,000 and $250,000, each fund four ways, against cash, against the gold sleeve as shipped, and against gold held only while the engine is out. Alone over their windows the funds earned 1.6% to 9.6% a year with worst falls of -20% to -36%, and moved with the Nasdaq fund at a correlation of -0.16 to +0.18: the diversification is real, the return is not. Beside the engine, none of the sixteen constructions met the pass mark on any dial. The timing is the autopsy: the trend funds’ 2022 was made in the first half of the year, before the exit turned off in June; inside the out-stretch that followed (2022-06-01 to 2023-04-03) they gave back -12.8% to -0.0% while gold made +7.1% and cash +2.7%. In the 2008 stretch the one fund old enough returned +1.1%, cash +1.4%. The gold sleeve pays because gold tends to rise in the stretches the engine sits out; trend funds tend to have already risen by then. Same lesson as the energy seat (grave 154): the thing that worked in 2022 worked before the exit knew.

our full test on the R-012 harness, v2 sleeve machinery, listing to 2026-09-11, $25,000 and $250,000, shipped rules · alone: RYMFX 1.6% / -36% / Sharpe 0.21 / corr QQQ +0.13 (from 2007-02-22) · AQMIX 4.6% / -26% / Sharpe 0.51 / corr QQQ -0.01 (from 2010-01-05) · DBMF 9.6% / -20% / Sharpe 0.80 / corr QQQ +0.18 (from 2019-05-08) · KMLM 7.4% / -31% / Sharpe 0.56 / corr QQQ -0.16 (from 2020-12-02) · best construction per fund on ULTRA, Sharpe minus gold, $25k / $250k: RYMFX -0.071 / -0.061, AQMIX -0.044 / -0.065, DBMF -0.033 / -0.085, KMLM -0.052 / +0.029 · best cell anywhere: KMLM trend-always ULTRA $250k +0.029 over gold (its other size fails) · engine-out 2022-23: RYMFX -5.6%, AQMIX -0.0%, DBMF -11.5%, KMLM -12.8%, gold +7.1%, cash +2.7% · docs/research/R-054-results.md
B-RANKmeasured

A sector all-weather book, alone and as a sleeve

A sector book is the market wearing a costume; it moves with the engine and pays less.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A four-bucket sector book from another chat, tested the way every sleeve is tested here: alone it earned 11.3% a year with a -30% worst fall, less than the safest dial with a deeper fall, and beside any dial it sat below the gold sleeve at both account sizes. A search for the best blend of the sector shelf, with the picking rule written down first, chose bonds, gold and utilities and failed out of sample too.

THE BOOK AS A 20% SLEEVE, MINUS THE GOLD SLEEVE, EACH DIAL-0.08ULTRA-0.05-0.07SELECT-0.07-0.06STEADY-0.06grey $25,000 · red $250,000 · gated construction
read the autopsy

Mel brought a “sector all-weather” book from another chat: growth sectors and semiconductors for good times, health care, staples and utilities for bad ones, energy and gold for inflation, long Treasuries for the opposite, with a five-year backtest that beat the S&P 500 and a bond-heavy 60/40 on every line. The five years were 2021 to 2026, one regime, and its two headline claims flipped on the full cycle: energy’s correlation with the market is 0.27 in that window and 0.60 since 2005, long Treasuries’ is 0.56 in the window and below zero since 2005. We wrote the rules and the pass mark down before the test ran. The book, without the memory-chip fund that listed in April, was run from 2005, which holds 2008, 2020 and 2022, three ways: alone against the three dials; as a fifth of the book beside each dial, held outright and on its own trend gate, at $25,000 and $250,000, against the gold sleeve; and as one of five candidate blends of the sector shelf, with the blend chosen on 2005 to 2015 and judged on 2016 to 2026 so nothing could be picked on the answer. Alone: 11.3% a year, worst fall -30%, against STEADY’s 13.6% and -18%; it lost 2008 by more than twice as much and moves with the engine (correlation 0.65 with ULTRA, 0.85 with the Nasdaq fund). As a sleeve: below the gold sleeve on every dial at both sizes, the best construction by -0.05 Sharpe, and every construction deepened the worst fall. The blend search: the only candidate with a positive training score was the four funds least like the engine, bonds, gold and utilities in equal weight, and on the test window it sat -0.07 below gold on ULTRA and failed every dial. The lesson is the one graves 15, 17, 150 and 151 already carry: sectors are the market in pieces, the pieces move together when it matters, and the one piece that does not, bonds, lost money in the stretch the engine was out.

our full-history test on the R-012 harness, 2005-01-03 to 2026-08-31 (GLD lists 2004-11), $25,000 and $250,000, shipped rules, DRAM excluded · alone: book 11.3% / -30% / Sharpe 0.96, IEF split 11.3% / -30%; STEADY 13.6% / -18% / 0.99, SELECT 16.4% / -20%, ULTRA 18.2% / -28% · 2008 book -17.5% vs STEADY -6.9%; 2022 book -10.2% vs STEADY -11.8% · sleeve Sharpe minus gold, gated, $25k / $250k: ULTRA -0.075 / -0.047, SELECT -0.065 / -0.065, STEADY -0.056 / -0.056 · blend search: training scores book -0.023, book_ief -0.054, equal_sectors -0.046, invvol_sectors -0.043, least_like_engine +0.022; selected least_like_engine, test ULTRA gated -0.073 · docs/research/R-053-results.md
B-RANKmeasured

Futures when violent, the two-times fund when calm

Switching between the two-times fund and futures by how violent the market is lost to plain futures in every era.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

For the futures version of the engine, we tried carrying the exposure in futures only when the market is violent and in the two-times fund when it is calm, at four settings and two account sizes. Every setting earned less than always using futures, and none of them won in all three eras.

SWITCHING SETTING MINUS ALWAYS-FUTURES, RETURN PER YEAR AFTER TAX-0.07-0.31above 0.15 vol-0.28-0.34above 0.2 vol-0.15-0.72above 0.25 vol-0.15-0.89above 0.3 volgrey $40,000 · red $100,000
read the autopsy

The futures version of the engine holds its Nasdaq exposure as micro futures contracts with the leftover in the plain and two-times funds; it was scoped in early September and is not sold. Its one known weak era is 2009 to 2016, a calm, steady bull, where a two-times fund that resets every day compounds better than futures that carry a fixed amount, while in choppy years the reverse holds. The engine already measures how violent the market has been over the last twenty sessions, so the natural idea was a switch: futures when the market is violent, the two-times fund when it is calm. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards. First the model itself was moved into the repository with its inputs frozen and made to reproduce its September figures exactly; then two honesty checks were run on it before any tuning: real day-by-day interest rates on the cash it holds instead of a flat two percent, which left the futures version’s edge unchanged within a tenth of a percent, and the broker’s own continuous futures prices for the two years it serves, which matched the model’s assumed cost of carry within a tenth of a percent a year. Then the switch, at four settings and two account sizes, judged against the plain futures version after tax over the whole history and in each of three eras. Every setting earned less over the full history, the best by -0.07% a year at $40,000, and none won all three eras: the two settings that helped in 2009 to 2016 lost in the other two, and at $40,000 they held no contracts at all on average, which makes them the plain-fund version under another name. The pass mark had three parts and this failed the first, so the other two were not needed. What stands: the plain futures version as scoped, unchanged, with its honesty checks now on the record. What is still outside the model, stated: margin calls, contract months before 2019, and an order path for futures in the sold software, which does not exist yet.

our full-history test, 1999 to 2026, after tax, fixed-ratio convention, ULTRA dial · step 0 reproduced the 2026-09-02 table exactly ($40k +1.02, $60k +1.15, $80k +1.95 at equal worst fall) · real daily rates: +1.08 / +1.26 / +1.81 · broker continuous NQ 2024-03-18 to 2026-09-14: realised excess 16.75% vs model 16.80% a year, daily correlation 0.9971 · the switch, return per year minus always-futures, $40k / $100k: above 0.15 vol -0.07 / -0.31 · above 0.2 vol -0.28 / -0.34 · above 0.25 vol -0.15 / -0.72 · above 0.3 vol -0.15 / -0.89 · era edges at $100k, above 0.20: -0.54 / +0.50 / -0.87 · gate: better in all three eras at both sizes, then an equal-fall control and a mechanism test; failed at the first · status: measured; the futures engine itself stays scoped, not sold
B-RANKmeasured

Energy, held only while the engine is out

Energy’s good year came before the Nasdaq exit turned off; inside every out period, gold beat energy.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We held the energy sector, and five other energy funds, only during the three stretches the engine’s Nasdaq exit was out since 1999, with cash the rest of the time, and compared that seat to cash, to gold, and to gold held the same way. Inside those stretches energy lost -29% and -29% and then made +2%, while gold made +30%, -2% and +7%; no construction met the pass mark on any dial.

WHILE THE NASDAQ EXIT WAS OUT: QQQ · CASH · GOLD · ENERGY (XLE), TOTAL RETURN-66%+7%+30%-29%2000 to 2003-14%+1%-2%-29%2008 to 2009+5%+3%+7%+2%2022 to 2023
read the autopsy

Energy was the one thing that rose in 2022 while the Nasdaq fell, so the natural question was whether the engine should hold energy in the months its Nasdaq exit has it out of the market. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards. The out-state is the engine’s own monthly reading of the Nasdaq against its 200-day average, which gives three stretches out: 2000-11-01 to 2003-05-01, 2008-03-03 to 2009-06-01, and 2022-06-01 to 2023-04-03; STEADY’s S&P leg has its own exit, so on STEADY this is the Nasdaq leg’s out-state, not the whole book’s. Six energy funds were held in the 20% seat only during those stretches, plain and with their own trend gate, on every dial at $25,000 and $250,000, against the same seat holding cash, holding gold, and holding gold only during the same stretches. The switching happens inside a stand-in return series for the seat, not as actual shares with fees and settlement, which is the usual simplification of these seat tests and is stated here. Inside the out periods: in the dot-com bear the energy sector lost -29% while gold made +30% and cash +7%. In 2008 the sector lost -29%; gold lost -2% and cash made +1%, so gold beat energy there but not cash. In 2022 the exit did not turn off until June, by which time the energy rally had already happened: from there to the re-entry in April 2023 the sector made +2%, the producers fund lost -14%, gold made +7% and cash +3%. The calendar year and the engine’s year are not the same thing, and the famous energy year falls in the part the engine was still in. Funds listed after a stretch began are measured from their listing inside it. No construction met the pass mark, which asks for a clear gain over cash, over gold, and over gold held the same way, at both sizes, with no deeper fall than gold. The midstream fund AMLP edged cash by a hair in several rows and never came near the gold seats; every other fund trailed all three. What this settles is narrow: these six energy constructions did not earn the seat. It does not settle energy’s long-term merit as an investment, and it does not close the watch on the young midstream income fund from the stone before this one. The book keeps what it has.

our full-history test, 1999 to 2026-09-11 · out stretches from the engine’s own Nasdaq exit: 2000-11-01 to 2003-05-01, 2008-03-03 to 2009-06-01, 2022-06-01 to 2023-04-03 · total return inside each stretch, QQQ / cash / gold / XLE: 2000: -66% / +7% / +30% / -29% · 2008: -14% / +1% / -2% / -29% · 2022: +5% / +3% / +7% / +2% · funds: XLE, IXC, VDE, XOP, AMLP, MLPX (IXC from its 2001 listing inside the first stretch; VDE, XOP, AMLP, MLPX absent from the stretches before their listing) · gate: return per unit of risk (Sharpe: a difference under 0.03 is nothing) at least 0.03 above cash-always, gold-always and gold-while-out at both sizes, no deeper fall than gold-always · AMLP plain on ULTRA at $25k: +0.005 over cash, -0.074 against the gold seat · the seat switches inside a stand-in series, not as executed shares · status: measured, corrected after the Lab’s review, awaiting reproduction
B-RANKmeasured

The NEOS income shelf as a sleeve

Most of these funds trail the index they are built on; the three that met the numbers are young and have only seen a rising market.
failed on its ownTested as a strategy on its own, after trading costs.

We carried seventeen NEOS income funds beside the engine, on every dial, at two account sizes, against the bills seat and the gold seat. Eleven trail the plain index or asset they are built on; none passed as a sleeve, and the three that met the numbers, SPYH, NLSI and MLPI, have only seen a rising market: SPYH has seventeen months of record, the other two under a year.

EACH FUND MINUS ITS OWN INDEX, RETURN PER YEAR, SINCE LISTINGQQQHSPYHMLPIIAUIXBCIXQQIBTCINIHIIWMIHYBIIYRINEHIXSPIBNDICSHITLTINLSIgreen: ahead of its reference · red: behind it · eleven of seventeen trail
read the autopsy

Most of these funds own an index, or a single asset like bitcoin or gold, and sell options on it every month; the option buyer pays cash now and takes most of the gains above a certain level. That cash is the distribution, so the fund hands you your own upside as income and keeps a price that barely moves. The long-short fund and the bond funds work differently and are read on their own terms. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards. Every fund on the issuer’s page except SPYI and QQQI, which were tested last week, was measured two ways: against a reference for what it holds over the identical window (the plain index for the equity funds; SPY as a broad reference for the long-short fund and a pipeline index for MLPI, neither a replica of the fund), and as a 20% sleeve beside the engine on every dial, held two ways (always in, or in only while above its own 200-day average), at $25,000 and $250,000, judged against the same seat holding bills and holding gold. Against its reference, eleven of seventeen trail, most by two to eight percent a year with nearly the same worst fall; six are ahead (NEHI, XSPI, BNDI, CSHI, TLTI, NLSI), the bond ones by about one percent a year. Of the four older funds these designs copy, the Nasdaq and S&P covered-call funds since 2013 and the gold one since 2013 trail their index by three to eleven percent a year; the Treasury one, since 2022, beat a Treasury index that was itself falling. Three of the four had a shallower worst fall than their index, so the family read is: less rise, and only sometimes less fall. Beside the engine, nothing passed. Three met the numbers on some dials: SPYH, a hedged S&P fund, on every dial; NLSI, a long-short fund, on every dial; MLPI, an energy-pipeline fund, on SELECT and STEADY. All three are young, SPYH at seventeen months and the other two under a year, and have only ever traded in a rising market, where a fund that sells calls looks its best; the rule is that such a window can earn a watch, never a pass. SPYH’s daily moves also track the engine’s (correlation 0.83) and the Nasdaq fund’s (0.93), which fails the different-asset bar whatever its numbers say. NLSI trades with a spread of about a quarter of a percent, which no test here charged. The rest: two of the three boosted funds, launched this February, are behind their index already and the third is level with it; the bitcoin income funds have lost 26% and 35% of their price since listing (their worst falls 37% and 48%), the ether one 37% of its price with a worst fall of 50%; the gold income fund trails gold by about eight percent a year with the same fall; the bond ones sit near zero and add nothing a bills seat does not. QQQH’s price history reaches back to 2019 because the ticker belonged to a predecessor fund whose option strategy was later changed, so its record through 2022 is the predecessor’s, not the current fund’s. What stays open: NLSI and MLPI hold their watch until their first real fall, and neither is a candidate for the book before then. The book keeps what it has.

our full-history test where the record allows, else since listing, to 2026-09-11 · 17 funds, 3 dials, 2 constructions, 2 sizes · fund minus its reference, return per year: QQQH -10.7% · SPYH -8.6% · MLPI -8.4% · IAUI -7.6% · XBCI -6.6% · XQQI -5.6% · BTCI -2.2% · NIHI -2.0% · IWMI -2.0% · HYBI -0.8% · IYRI -0.0% · NEHI +0.1% · XSPI +0.3% · BNDI +1.0% · CSHI +1.0% · TLTI +1.1% · NLSI +11.1% · families, fund minus index per year and worst fall vs index: QYLD -10.6% (-25% vs -35%), XYLD -6.2% (-34% vs -34%), TLTW +3.0% (-19% vs -24%), GLDI -2.8% (-34% vs -38%) · gate: return per unit of risk (Sharpe: a difference under 0.03 is nothing) at least 0.03 above both the bills seat and the gold seat at both sizes, no deeper fall than gold (STEADY: unchanged STEADY, one comparison) · watch: SPYH (every dial, both constructions; moves with the engine 0.83, with QQQ 0.93), NLSI (every dial, fixed, 189 sessions), MLPI (SELECT and STEADY, fixed, 183 sessions) · different-asset bar as registered = moves with the engine at most 0.45 (as computed in the run: with QQQ at most 0.45 and gold at most 0.5; the two readings agree on every fund here); cleared by QQQH, NLSI, IYRI, MLPI, CSHI, BNDI, TLTI · QQQH history includes a predecessor fund with a changed option strategy · corrected after the Lab’s reproduction (numbers matched to 1e-13; eleven not fourteen trail; family and crypto figures restated) · status: measured, reproduced
B-RANKmeasured

The best blend of the survivors as a sleeve

Blending the funds did not improve the book, and this way of testing blends cannot settle the question.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We let five recipes build the best mix of twelve funds, gold and bills, each January using only the past, and judged the years after. No mix improved the book beside gold. A result that briefly looked like bills beating STEADY’s gold turned out to be a flaw in how the mix was fed to the engine: tested directly, bills lose to gold there.

EACH BLEND AS THE SLEEVE, 2019 TO 2026, SHARPE MINUS THE GOLD SLEEVE-0.067equal-0.128minvar-0.092maxsharpe-0.064riskparity-0.005lowcorr
read the autopsy

We took every fund from this weekend that had beaten a bills sleeve on at least one dial, twelve of them plus gold and bills, and let five different recipes build a blend from them: equal weight, lowest volatility, highest return per unit of risk, equal risk from each part, and one that directly minimises how much the blend moves with the engine. Each recipe could only look backwards, re-picking its weights every January from the years before, using only funds already listed by then, and only 2019 to 2026 was judged, which holds the 2020 crash and the 2022 bear. The mix was held as whole shares at real prices from the sleeve’s own dollars, cash in bills, monthly rebalances, next-day settlement and a dollar per fund traded, and then handed to the engine as a single price series standing in for the sleeve. This stone went through four passes. The Lab found five defects in the first run, the reviewer three in the second, and the third pass answered them; the Lab then tested the one thing the third pass seemed to show, that a nearly-all-bills sleeve beat STEADY’s gold, by putting bills in place of gold directly inside the engine: bills lose there, two points a year of return and about a tenth of Sharpe at both sizes, with a slightly shallower fall. So that reading was an artefact of the stand-in series, and it is withdrawn. Three accounting gaps in the method remain and are why it stops here rather than being patched again: the sleeve’s own cash and the book’s cash are kept separately; the mix sizes and fills its constituents at the same close, which is the timing flaw fixed elsewhere in the engine; and the least-correlated recipe treated a fund’s pre-listing days as zero returns when measuring correlation. Settling the blend question properly needs the constituents inside the engine as real lines, which is a feature the engine does not have and would have to be built and audited before any blend could pass. What can be said, and only this: on every dial and at both sizes, none of these five recipes improved the book over the gold sleeve, and the covariance estimates two of them rest on were unreliable (most yearly fits needed repair). It does not establish that no blend can, and it does not establish that only gold can. The book keeps what it has.

four passes; walk-forward yearly fits from 2016; sleeve as real shares at real prices, judged 2019-01 to 2026-08; every recipe below the gold sleeve on ULTRA and SELECT at $25,000 and $250,000 (least-correlated recipe, the closest: about −0.03 of Sharpe vs gold on ULTRA at $250,000) · the STEADY near-cash ‘passes’ of the third pass withdrawn: bills placed directly in STEADY’s gold seat inside the engine read −1.98 / −2.36% a year and −0.085 / −0.100 Sharpe vs unchanged STEADY at $25k / $250k, worst fall shallower by 0.7 / 0.6 points (the Lab’s direct test) · covariance: 130 of 190 yearly fits were not valid covariance matrices and were repaired · remaining method gaps: separate inner and outer cash, same-close sizing and fills for constituents, zero-filled pre-listing returns in the correlation objective · status: closed as not settled by this method; reproduced by the Lab through the third pass
B-RANKmeasured

The rest of the issuer’s shelf: bonds, credit and hybrids as a sleeve

The bond and hybrid shelf failed the sleeve test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We ran the nineteen iShares funds without a stone, bonds, credit, mortgages, munis, REITs and a nine-month alternatives fund, the same way. Bonds were the only ones that move against the engine, and they lost money in the one stretch it was out.

BEST OF SIX CONSTRUCTIONS PER FUND: SHARPE MINUS THE GOLD SLEEVE (GATE +0.03)+0.00USMV-0.00TIP-0.00TLT-0.01IGSB-0.01FLOT-0.01IEF-0.02GOVT-0.02MUB-0.03MBB-0.03LQD-0.03IGF-0.03ICVT
read the autopsy

Mel sent the issuer’s whole shelf, five hundred and twenty-four products sorted by assets, and asked which could be sleeves. Most of it already had a stone: broad and factor equity (17, 19, 117, 149), sectors and themes (15, 16, 121, 150), countries (150), dividend funds (149), preferreds (7), option-income and buffer-style funds (146, 147), TIPS and the short Treasury funds (one of which is the book’s cash), gold and silver (the incumbent sleeve), bitcoin (its own gated seat), ether (worse than bitcoin, September 2nd), semiconductors (16, 121). What had never been run as a sleeve was the fixed-income and hybrid middle of the lineup: mortgages, investment-grade and short corporate credit, two high-yield funds, emerging-market bonds, municipals, convertibles, floating rate, the Treasury curve at three maturities, TIPS again, the flexible-income active fund, minimum volatility, global infrastructure, two REIT funds, and the systematic-alternatives fund launched last December. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards on the same design as the country and theme stone: each a fifth of the book, held outright or on its own 200-day gate, on all three dials, against each dial’s bills and gold sleeves, corrected runner, frozen inputs. 114 constructions. None passes. The bond funds are the honest surprise for anyone who expects bonds to diversify a stock engine: the Treasury curve is the only thing on the shelf with a negative correlation to the engine (-0.13 for the long end), and it earned about nothing over the window (0.1% a year with a 48% fall) and lost 7% in the ten months of 2022 the engine was out, the one stretch it was held for. Credit and high yield carry equity risk (0.50, 0.46) and pay a bond’s return for it. Mortgages and municipals are near-zero correlation and near-zero return, which is what bills are, with a fall. Convertibles are equity in a costume (0.64). Minimum volatility beats bills gated on ULTRA by +0.015 and gold by +0.005, the best number on the shelf and a sixth of the gate. One line is not buried: the systematic-alternatives fund, a multi-strategy book of trend, carry and the like with no equity leg, is nine months old, has never seen a bear, and reads 23% a year with a 2% fall over that stretch; every construction of it clears the gate by a margin that means nothing on nine months. By the registered rule it was a watch, the only one on the shelf, until the Lab found the same strategy family with ten years of record in its mutual-fund siblings and we ran them through the same gate: correlation with the engine 0.17 and 0.12, 7.3% and 4.9% a year with falls under ten percent, and as a sleeve the better of them beats bills on SELECT by +0.056 and loses to the gold sleeve on every dial (-0.068 on SELECT, -0.088 on ULTRA). A low-correlation stream that earns less than gold is a slower bill with a fee. The watch is closed on its family’s record, with the Lab’s caveat kept that these are the siblings’ histories, not the new fund’s. The shelf is otherwise closed: on this book, a bond is a slower bill with a drawdown, and a hybrid is equity with a story.

corrected runner, frozen inputs, 19 funds x 2 constructions x 3 dials = 114 constructions, own window (listing or 2013) to 2026-08 · correlation with the engine: TLT -0.13, IEF -0.13, GOVT -0.12, TIP -0.01, MBB 0.05, MUB 0.06, LQD 0.15, FLOT 0.15, EMB 0.36, HYG 0.50, USHY 0.46, ICVT 0.64, USMV 0.60, IGF 0.49, REET 0.42 · best Sharpe over the gold sleeve, any dial or construction: USMV +0.005 · TIP -0.003 · TLT -0.005 · IGSB -0.009 · FLOT -0.012 · IEF -0.013 · GOVT -0.021 · MUB -0.021 · alone since 2013: TLT 0.1%/yr, fall 48%; IEF 1.1% / 24%; HYG 4.2% / 22%; MBB 1.5% / 18% · while the engine was out in 2022: TLT -7.2%, IEF -0.9%, HYG +1.1%, REET -13.5% · IALT since 2025-12: watch closed by a second test — siblings BDMIX / BIMBX since 2016: corr 0.17 / 0.12, 7.3% / 4.9%/yr, falls 10% / 9%; as a sleeve vs bills / gold: BDMIX SELECT +0.056 / -0.068, ULTRA +0.003 / -0.088
B-RANKmeasured

Emerging markets, robots, quantum and the rest, fixed and gated, on every dial

Themes and countries failed the sleeve test, gated or not.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We ran twenty-seven emerging-market, robotics, quantum, AI and other funds as a fifth of the book, held outright and on their own trend gate, on every dial. They all move with the engine; none beat gold.

BEST OF SIX CONSTRUCTIONS PER FUND: SHARPE MINUS THE GOLD SLEEVE (GATE +0.03)+0.02XLE+0.01SMH-0.01ITA-0.02IHI-0.02EWY-0.02QTUM-0.03EWT-0.05ROBO-0.05EFA-0.06MCHI-0.06LIT-0.06TAN
read the autopsy

Mel: emerging markets, robots, quantum, and so on. The record already said what a fixed slice of these would do (grave 17, higher-volatility assets generally; 15, sector rotation; 16 and 121, semiconductors; 19, spreading the rule across markets; the 44-ETF expansion screen of September 3rd, where every large edge turned out to be the same asset as the root), but one construction had not been run on them: the gated sleeve, held only while the fund is above its own 200-day average at the month’s first session and in bills otherwise, which is how bitcoin and gold are held. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards: twenty-seven funds, ten geographic (emerging markets broad, India, China, China internet, Brazil, Japan, Korea, Taiwan, developed ex-US) and seventeen thematic (robotics twice, quantum, AI, innovation, space twice, clean energy twice, uranium, lithium, biotech, medical devices, defence, infrastructure, energy, semiconductors), each as a fifth of the book both fixed and gated, on ULTRA, SELECT and STEADY, against each dial’s bills and gold sleeves, corrected runner, frozen inputs. 162 constructions. None passes. The correlations say why before the simulations do: every one of these funds moves with the engine’s own returns, from 0.32 for energy to 0.81 for the AI fund, and the themes cluster around 0.72 to 0.76; they are the root with more volatility and a story. The gate lets a few of them beat bills, quantum and semiconductors and infrastructure on ULTRA, energy on SELECT (+0.048); none beats the gold sleeve, the best being energy gated on SELECT at +0.003 and semiconductors fixed on ULTRA at +0.001. The gate helps the worst of them (China, uranium, innovation) by taking them out of their falls and cannot make them a sleeve. The emerging-market funds are the weakest of all: 5.2% a year alone for the broad fund since 2013 with a 40% fall, and negative on every dial, gated or not. The lesson is the same one the income and dividend stones taught, now on the growth side: a sleeve earns its place by being uncorrelated with the engine, not by being exciting, and every theme is correlated. What robotics, quantum and AI do have is a claim on the root itself, the semiconductor root (grave 121) and this one; the engine already holds the Nasdaq-100, which is where those themes live when they work.

corrected runner, frozen inputs, 27 funds x 2 constructions x 3 dials = 162 constructions, each fund’s own window (listing or 2013) to 2026-08 · correlation with the engine: EEM 0.65, EFA 0.65, BOTZ 0.72, ROBO 0.72, QTUM 0.76, AIQ 0.81, ARKK 0.65, SMH 0.76, XLE 0.32, URA 0.42 · best Sharpe over the gold sleeve, any dial or construction: XLE +0.017, SMH +0.014, QTUM -0.020, IHI -0.016, ITA -0.009, PAVE -0.088, EWT -0.028, EEM -0.085, ARKK -0.128, KWEB -0.163 · alone since 2013: EEM 5.2%/yr fall 40%, MCHI 2.6% / 63%, KWEB 1.8% / 81%, ARKK 13.9% / 81%, SMH 30.6% / 45% · while the engine was out in 2022: ARKK -19.5%, EEM -4.5%, XLE +13.2%
B-RANKmeasured

The dividend sleeve, hunted on all three dials

Dividend sleeves failed on all three dials.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We ran midstream payers, five best-method dividend funds and an active dividend-value fund on ULTRA, SELECT and STEADY. None improved the book, and replacing gold on STEADY made the falls deeper.

ULTRA + 20% SLEEVE: SHARPE MINUS THE GOLD SLEEVE, SAME WINDOW-0.052LVHD-0.049GCOW-0.106SPHD-0.081DGRW-0.081WDIV-0.084ET-0.046MLPX
read the autopsy

Mel had a feeling there was a dividend sleeve somewhere, and asked the Lab to scour on its own while the option-income shelf was being buried. The Lab took it in three passes on the corrected sleeve runner, each frozen before it ran, each reproduced independently. First the midstream names that pay the fattest distributions, an energy partnership and the pipeline fund: beside ULTRA over 2019 to 2026 they returned about what the gold sleeve returned and fell eight points deeper (-30% and -29% against -21%). Then the five dividend funds with the best-reasoned methods on the shelf: sustainable dividends with low price and earnings volatility (LVHD), global free-cash-flow screens (GCOW), high dividend with a low-volatility screen (SPHD), quality dividend growth (DGRW), global aristocrats (WDIV), each as a fifth of the book beside ULTRA over 2017 to 2026 against bills, gold, the S&P and the world index: every one had a lower Sharpe than both bills (1.158) and gold (1.194), the best of them (GCOW) at 1.145, and every one fell deeper than bills. Mel then asked the right question: was ULTRA the wrong host? So the same five ran on SELECT, with its lower exposure, and on STEADY, replacing its gold. On SELECT the best of them (LVHD) added +0.015 of Sharpe to the unchanged engine, below the declared benefit, with a lower return and a deeper fall, while the gold sleeve on the same book added +0.108. On STEADY every dividend replacement lowered Sharpe and deepened the worst fall, from 18.4% to between 22.6% and 25.2%. Ten more constructions, ten failures. The active dividend-value fund that had met the gate on a single-bear window (stone 147) was closed the same day: at ten and thirty percent it fails, and a block bootstrap of its twenty-percent result crosses zero against both the S&P and bills. None of this says dividend investing fails. It says a fixed slice of a dividend fund does not improve this book on any of its three dials, for the same reason the income funds did not: it carries the equity risk the engine already sizes, and the two sleeves that do not, bills and gold, are the comparison it has to beat. The feeling was worth testing. It is now tested.

the Lab’s dividend family, corrected sleeve runner, frozen inputs, independently reproduced · ULTRA + 20% sleeve from January 2017 to August 2026 (the funds’ own listings, not the full cycle; it holds 2020 and 2022 and no dot-com bear), Sharpe: LVHD 1.141, GCOW 1.145, SPHD 1.088, DGRW 1.113, WDIV 1.113 vs bills 1.158, gold 1.194 · midstream from January 2019 to August 2026 (the partnership’s clean record; not the full cycle): ET 1.029 (fall -30%), MLPX 1.067 (-29%) vs gold 1.113 (-21%) · SELECT + sleeve vs unchanged 1.149: LVHD 1.164, DGRW 1.159, GCOW 1.153, gold 1.257 · STEADY with the gold replaced, vs unchanged 1.184 / fall 18.4%: DGRW 1.108 / 22.6%, LVHD 1.104 / 23.7%, WDIV 1.075 / 25.2% · CGDV at 20%: bootstrap 95% interval vs S&P [-0.071, +0.153], vs bills [-0.128, +0.182]; 10% and 30% fail · distribution accounting approximate (adjusted prices), stated by the Lab
B-RANKmeasured

The uncorrelated sleeve, hunted directly

The uncorrelated-sleeve hunt failed against gold.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We measured what moves least with the engine, then ran the candidates with a return: trend futures, a convex tail fund, the dollar, and blends with gold. Each lost to the gold sleeve the book already has.

20% SLEEVE BESIDE ULTRA: SHARPE MINUS THE GOLD SLEEVE, SAME WINDOW-0.058KMLM-0.080CTA-0.059CAOS-0.045UUP-0.048GOLD+KMLM-0.082GOLD+CAOS
read the autopsy

Mel read the two income stones correctly: the only sleeve that has ever earned its place, gold, won because it is uncorrelated with the equity the engine already holds, so the riddle is to find the most uncorrelated thing with a return. Measured first, not guessed: every candidate’s correlation with the engine’s own daily returns on the full-history test, not with the index. The zeros with a positive record were bills and gold (both already in the book), pure trend futures with no equity leg, the dollar, a convex tail fund that has not bled, and the rising-rates fund already withdrawn; everything with a strongly negative correlation (tail funds, VIX products, anti-beta) pays for it in carry every year it is not needed, and Treasuries were a zero until 2022 and then were not. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards: each candidate and two gold blends as a fixed 20% sleeve beside ULTRA on the corrected runner, over its own window, against the same 20% in bills and, this time, in gold, because a new uncorrelated sleeve has to beat the one the book has; plus the sleeve must be positive on the engine’s worst days and its own failure mode must be in its window. Every candidate lost to the gold sleeve: pure trend -0.058 of Sharpe, trend-plus-carry -0.080, the convex fund -0.059, the dollar -0.045, half gold and half trend -0.048. Trend beat bills and lost 10% in the ten months of 2022 the engine was out, the one stretch a diversifier is for, because trend reversed as the engine re-entered. The convex fund is the interesting one: positive on the engine’s worst days, a fall of a few percent, but three years old with no calm year in its record, the year that decays it; watched, not passed. The honest footnote is about the incumbent: over this same 2013-to-2026 window the gold sleeve beats the bills sleeve by only +0.011 of Sharpe on ULTRA. Gold earned its seat on the long record (1975 on, with its own twenty-year bear) at equal drawdown; on the last thirteen years it is a wash against cash beside this engine. The bar for the next sleeve is therefore not high, and nothing cleared it. The most uncorrelated thing with a return that we can buy is the thing we already hold.

corrected runner, inputs archived, ULTRA 80 + 20% sleeve, each candidate’s own window to 2026-08 · Sharpe over the bills / gold sleeve: KMLM +0.039 / -0.058 · CTA +0.024 / -0.080 · CAOS +0.023 / -0.059 · UUP -0.034 / -0.045 · gold+KMLM +0.050 / -0.048 · gold+CAOS -0.000 / -0.082 · correlation with the engine: KMLM -0.03, CTA -0.07, CAOS -0.04, UUP -0.04, gold +0.04 · on the engine’s worst 5% of days: CAOS +9 bp, KMLM +0, CTA -8 · while out 2022-05 to 2023-03: KMLM -10.1%, CTA -5.3%, gold +5.5%, bills +2.4% · gold sleeve vs bills sleeve on ULTRA since 2013: +0.011 · calm years in the record (Nasdaq vol under 13%): [2013, 2014, 2017]
B-RANKmeasured

The JEPI tier, monthly income funds and a model portfolio’s parts as a sleeve

The JEPI tier and a model portfolio's parts failed the sleeve test.
failed on its ownTested as a strategy on its own, after trading costs.

We ran thirty-two monthly income funds and portfolio components as a fifth of the book against a fifth in bills and in gold. Beside a volatility-sized stock engine, anything with stock risk lost to bills.

20% SLEEVE BESIDE ULTRA: SHARPE MINUS THE BETTER OF ITS UNDERLYING AND BILLS+0.02JEPI+0.02JEPQ-0.11GPIX-0.08GPIQ-0.04SPYI+0.01OVL+0.01DIVO+0.01SPMO+0.00AVUV-0.01GDE+0.05DBMF+0.05CEFS+0.07RISR+0.03CGDV
read the autopsy

Mel’s follow-up: never mind weekly, what about the JEPI tier, monthly funds, and the components of a model portfolio making the rounds (leveraged tech, efficient gold, gold miners, momentum, managed futures, commodities, small value, international value, a closed-end-fund fund, a rising-rates hedge, income funds)? We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards on a test whose rules were written down in advance’s design: thirty-two funds with a usable record, each against its own asset class over the identical window, then as a fixed 20% sleeve beside ULTRA on the full-history test against the same sleeve in the asset class and in bills, then the 2022 stretch the engine was out, then the tax arithmetic. This stone was regenerated after the Lab’s audit found four defects in the first runner (the sleeve book had run on SELECT’s settings; a sleeve whose underlying was the Nasdaq had pinned the core’s own signal; gates compared rounded figures; inputs were not frozen). The corrected runner fixes all four and the inputs are now archived; the numbers here are the corrected ones. The record is more mixed than the weekly shelf: the option-income funds still trail what they hold (JEPI -6.8 points a year, JEPQ -4.3, GPIX -2.6), while several of the portfolio’s components beat theirs (S&P momentum +4.2 over a decade, small value +4.9, international value +5.2, the gold funds +8.1 and +8.2 over gold and the miners). As a sleeve almost none of that survives, and the reason is the book it sits beside: ULTRA is already a volatility-sized equity position, so a fifth of anything with equity beta adds correlated risk the dial then has to cut, and a bills sleeve is hard to beat. Momentum beat its index by +0.031 and bills by +0.014, under the gate. Small value, international value, the gold funds, managed futures, the closed-end-fund fund: the same shape, ahead of their asset class, not ahead of bills by enough, or with a deeper fall. JEPI beats bills by +0.019; the rising-rates fund RISR beats them by +0.068 but with a deeper fall, on a window that is one regime. One fund met the gate: CGDV, an active dividend-value fund, +0.045 of Sharpe over the S&P sleeve and +0.034 over bills with an equal fall, on a window from February 2022 that contains one bear, the 2022 one, in which value beat growth by the widest margin in twenty years. That is the era a dividend-value fund is built for, and dividend-and-value ballast has a stone of its own (117). It went to the Lab with that caveat and came back closed the same day (stone 149): at ten and thirty percent it fails, and a block bootstrap of the twenty-percent result crosses zero against both comparators. While the engine was out in 2022, nothing on the list was a cash substitute; the managed-futures fund, sold as the crisis hedge, lost 12% in those ten months. A Roth removes tax drag; it changes no verdict.

thirty-two funds with a usable record (CAGE none), Yahoo total return to 2026-09, inputs archived · 20% sleeve beside ULTRA, Sharpe fund / asset class / bills: CGDV 1.21 / 1.17 / 1.18 (gate met on a one-bear window, to the Lab) · JEPI 1.14 / 1.1 / 1.12 · RISR 1.09 / bills 1.02 (fall -21% vs -20%) · SPMO 1.06 / 1.03 / 1.05 · AVUV 1.14 / 1.12 / 1.13 · GDE 1.17 / gold 1.19 / 1.06 · DBMF 1.11 / bills 1.06 · while out 2022-05 to 2023-03: bills +2.4%, CGDV +5.6% (worst -17%), JEPI +3.9% (worst -10%), DBMF -12.3% (worst -20%) · bull-window-only watches: QQQI, IDVO, ORR
B-RANKmeasured

The weekly-pay option-income funds as a sleeve

Weekly-pay option-income funds failed the sleeve test.
failed on its ownTested as a strategy on its own, after trading costs.

We ran twenty-three of them against the fund each one wraps, then as a fifth of the book, then in a Roth. They trail what they hold with the same falls; a Roth removes the tax on them and not the upside they sold.

TOTAL RETURN MINUS ITS OWN UNDERLYING, PP A YEAR, EACH FUND’S WINDOW-11QYLD-6XYLD-7JEPI-4JEPQ-1QDTE-2XDTE-11YMAX-19ULTY+3IDVO
read the autopsy

Mel’s question, and a fair one: the weekly-pay option-income funds from Roundhill, YieldMax, REX and the rest are the fastest-growing shelf in retail; is there a sleeve in the best of them, and does a Roth change the answer? We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards with the whole shelf frozen: twenty-three funds, each against its own underlying over the identical window, then as a fixed 20% sleeve beside ULTRA on the full-history test against the same sleeve in the underlying and in bills, then the one episode the engine was actually out that any of them lived through (May 2022 to March 2023), then the tax arithmetic. This stone was regenerated after the Lab’s audit found four defects in the first runner (the sleeve book had run on SELECT’s settings; a sleeve whose underlying was the Nasdaq had pinned the core’s own signal; gates compared rounded figures; inputs were not frozen). The corrected runner fixes all four and the inputs are now archived; the numbers here are the corrected ones. The record: with every distribution reinvested, twenty-two of the twenty-three trail the thing they hold, by 11 points a year for QYLD over twelve years and 19 for ULTY, with the same worst falls; up-capture equals down-capture, which is what selling the upside for premium has to produce. The one exception, IDVO, holds a different portfolio from its benchmark, and the semiconductor fund’s 111% a year is the semiconductor rally minus fourteen points. As a sleeve, nothing passes: JEPI, the best of them, beats the S&P sleeve by +0.037 of Sharpe and the bills sleeve by only +0.019, under the gate; the four that beat both did so on windows with no bear in them and are watched, not passed. While the engine was out in 2022, none was a cash substitute: bills made +2.4% with no fall, the funds -2.2% to +3.9% with falls of ten to twenty percent. And the Roth: it removes the tax drag, which is enormous on the weekly payers (15 points a year on QDTE, 26 on ULTY at a worst-case ordinary rate), and it cannot remove the upside the fund has sold, which is the gap to the underlying that remains in a Roth (-1.4 QDTE, -10.8 YMAX, -10.6 QYLD). A Roth makes them less bad. It does not make them a sleeve.

twenty-three funds, Yahoo total return to 2026-09, inputs archived · own-window gap to the underlying: QYLD -10.6%, XYLD -6.2, JEPI -6.8, JEPQ -4.3, QDTE -1.4, XDTE -2.3, YMAX -10.8, ULTY -18.9, CHPY vs SMH -13.9 · 20% sleeve beside ULTRA, Sharpe fund / underlying / bills: JEPI 1.14 / 1.1 / 1.12 (fail), QYLD 0.92 / 0.94 / 0.99, QDTE 1.03 / 0.93 / 0.92 (bull window, watched) · while out 2022-05 to 2023-03: bills +2.4%, JEPI +3.9% (worst -10%), QYLD -2.2% (worst -20%) · worst-case tax drag: QDTE +14.7% a year, ULTY +25.6, JEPI +3.3
B-RANKmeasured

Gold only when Treasuries also trend

A Treasury check on the gold sleeve failed the difference test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We held gold only when Treasuries were also rising, on top of the gold rule the book already has. The months it would drop were already months the gold rule sits out, so almost nothing changed.

STEADY SHARPE, 2003-2026, VS SHIPPED+1.015shipped+1.029with the check
read the autopsy

Pre-registered with the gate fixed in advance: at least +0.03 Sharpe on the STEADY book, no worse drawdown, two of three sub-windows, and a passing mechanism check on both roots; every variant enters the full-history test through the signal input only, so sizing, band, exit confirmation and execution are the shipped ones. A 2025 note on gold since 1969 reports that gold’s momentum is reliable only when Treasuries are also rising. On STEADY’s gold sleeve, which already holds gold only above its own 200-day average with a two-month confirmation, the extra check runs through that same confirmation (it does not exit the instant Treasuries turn negative), and it is nearly empty: the gold months it drops did average less (0.38% against 1.24%), so the mechanism is there, but nearly all of those months were months the sleeve was already out. Sharpe 1.015 to 1.029, worst fall unchanged, the same number of trades. Below the quarter of the gate, on a window that begins with the Treasury fund’s listing in 2002. Not added.

full-history test, STEADY, 2003-07 to 2026-08 (IEF from listing), $25,000 · Sharpe 1.015 shipped vs 1.029 with the check (+0.014) · worst fall -0.1% · CAGR 13.90% vs 14.09% · gold months kept / dropped by the Treasury check: +1.24% / +0.38%
B-RANKmeasured

A three-horizon vote as the exit

A three-signal exit vote failed the return test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We replaced the single trend rule with three horizons voting. The vote picked bad months correctly and still lost, because it sat out ordinary dips the single rule rides through.

SHARPE VS THE SHIPPED EXIT-0.136SELECT-0.119ULTRA-0.028S&P root
read the autopsy

Pre-registered with the gate fixed in advance: at least +0.03 Sharpe on both the Nasdaq and the S&P root, no worse drawdown, two of three sub-windows, and a passing mechanism check on both roots; every variant enters the full-history test through the signal input only, so sizing, band, exit confirmation and execution are the shipped ones. The trend literature says a vote across horizons of volatility-normalised returns beats a single moving average with less churn. The raw month-end vote does what it claims: the months it would drop averaged 0.44% on the Nasdaq against 1.33% for the months it kept, and the same on the S&P, so the mechanism check passes on both roots. That check counts the vote’s own months, not the smaller set the confirmed rule actually sat out, which the Lab’s audit noted. The book still loses, and badly: SELECT -0.136, ULTRA -0.119, with deeper falls, because a vote of three slow signals is out through long stretches of ordinary drawdown that the single rule with its two-month confirmation rides through, and the dial has already shrunk the position by then. Dropping bad months is not enough when the price of being out is paid in the good ones. The single 200-day rule with confirmation stays.

full-history test, 1999–2026, $25,000, shipped execution · Sharpe vs shipped: SELECT -0.136 · ULTRA -0.119 · S&P root -0.028 · worst fall vs shipped: SELECT -3.4% · S&P -0.4% · kept vs dropped months: Nasdaq +1.33% / +0.44%, S&P +1.17% / +0.26% (mechanism passes on the raw vote) · ULTRA worst fall -2.1%
B-RANKmeasured

Sizing on downside volatility only

Downside-only volatility failed the sizing test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We sized the position from the down days alone instead of all days. The measure was noisier, traded a fifth more, and the months it flagged were followed by better returns, not worse.

SHARPE VS THE SHIPPED DIAL-0.038SELECT-0.001ULTRA-0.017S&P root
read the autopsy

Pre-registered with the gate fixed in advance: at least +0.03 Sharpe on both the Nasdaq and the S&P root, no worse drawdown, two of three sub-windows, and a passing mechanism check on both roots; every variant enters the full-history test through the signal input only, so sizing, band, exit confirmation and execution are the shipped ones. The published claim (US factors over nine decades) is that the volatility of down days alone predicts next-month returns while total volatility does not, so scaling by it times the market as well as sizing it. On our roots the premise fails before the rule does: months whose downside volatility exceeded total volatility were followed by 1.28% on the Nasdaq and 1.27% on the S&P, higher, not lower, than the other months. The estimator as run is the root-mean-square of the negative days over the same twenty sessions, scaled by the square root of two so it equals total volatility on symmetric returns (the registration’s wording left that denominator ambiguous; the code is the rule). Built from half the days, it is noisier, and our full-history test paid for that in trades: 32 a year against 27. Sharpe fell on all three books. The twenty-day standard deviation stays.

full-history test, 1999–2026, $25,000, shipped execution · Sharpe vs shipped: SELECT -0.038 · ULTRA -0.001 · S&P root -0.017 · worst fall vs shipped: SELECT -3.2% · S&P +2.4% · trades a year SELECT 32 vs 27 · next month after downside > total vol: Nasdaq +1.28% vs +0.91%, S&P +1.27% vs +0.71%
B-RANKmeasured

Scaling only in the volatility extremes

Extremes-only sizing failed the two-root test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We sized the position only when volatility was unusually high or low and held a plain position otherwise, on the Nasdaq and the S&P. It helped one and hurt the other, and the reason it was supposed to work showed up on only one of them.

SHARPE VS THE SHIPPED DIAL-0.046SELECT-0.018ULTRA+0.032S&P root
read the autopsy

Pre-registered with the gate fixed in advance: at least +0.03 Sharpe on both the Nasdaq and the S&P root, no worse drawdown, two of three sub-windows, and a passing mechanism check on both roots; every variant enters the full-history test through the signal input only, so sizing, band, exit confirmation and execution are the shipped ones. The published result (ten equity index futures over four decades) says continuous inverse-volatility scaling over-trades in ordinary months, and that only the extremes carry information. On our book it split the roots: the S&P version improved by +0.032 with a shallower worst fall, the Nasdaq version lost 0.046 and its worst fall deepened by 6.3 points, because the months the rule leaves unscaled are exactly where the Nasdaq’s fast vol changes live. The mechanism check agreed with the split: after top-fifth-volatility months the Nasdaq averaged 0.10% a month against 1.39% in the middle, but the S&P averaged 1.22% against 0.82%, the wrong way round. A rule that works on one root and not the other is a fit, as the wide-band lesson taught. The continuous dial stays.

full-history test, 1999–2026, $25,000, shipped execution · Sharpe vs shipped: SELECT -0.046 · ULTRA -0.018 · S&P root +0.032 · worst fall vs shipped: SELECT -6.3% · S&P +1.8% · next-month return after top-fifth vol / middle: Nasdaq +0.10% / +1.39%, S&P +1.22% / +0.82%
B-RANKmeasured

The two-times leg on a margin loan instead of the fund

Margin borrowing failed the small-account cost test.
small account or this brokerTested as a small account paying the broker's margin rate. Bigger accounts borrow cheaper.

We compared buying the plain fund with borrowed money against using the leveraged fund. Under the tested broker rates, borrowing cost more at smaller account sizes.

MARGIN LOAN MINUS QLD, ULTRA, CAGR (PP/YR) BY ACCOUNT TIER-0.1%<100k-0.1%100k-1M+0.1%1M-3M-0.7%Lite
read the autopsy

The idea: instead of holding the two-times fund, hold the plain fund and borrow the extra from the broker. You would swap the fund’s fees for the broker’s interest, and the question is which costs less. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards. The two-times fund charges a yearly fee (0.95%) plus the cost of the borrowing it does internally, which runs a little above the short-term interest rate. Borrowing from the broker costs the same short-term rate plus the broker’s markup, which shrinks as the account grows: 1.5% under $100,000, 1% up to a million, 0.75% up to three million, and 2.5% on the free-tier plan; plus you pay the plain fund’s small fee on twice the money. Everything else in the engine stayed the same, so only the cost of leverage changed. Result: borrowing loses under $100,000 (0.14% a year worse on ULTRA) and badly on the free tier (0.67% worse); it is a wash from $100,000 to a million; it edges ahead only between one and three million, by 0.07% a year, a third of what we would need to see to change anything. Two things the test does not include, stated plainly: margin interest can be tax-deductible in a taxable account and fund fees cannot, and the two ways of holding leverage drift slightly differently between rebalances. Neither is big enough to close a gap this size for a small account. The two-times fund stays.

our full-history test, 1999 to 2026, short-term rate averaging 2.14% · yearly cost of each borrowed dollar, broker minus fund: under $100k +0.6% · $100k to $1M +0.1% · $1M to $3M −0.15% · free tier +1.6% · ULTRA yearly return, borrowing minus fund: under $100k −0.14% · $100k to $1M −0.05% · $1M to $3M +0.07% · free tier −0.67% · SELECT: −0.15% / −0.01% / +0.03% / −0.38% · the drift between the two ways of holding leverage is not modelled
B-RANKmeasured

Tax-loss harvesting on the engine’s own lots

Tax-loss harvesting failed the after-tax test on the engine.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We replayed the engine's own trades with lots and taxes, then sold losing lots to bank the loss and bought back the exposure. The engine's cuts had already realised nearly all of those losses, so the gain was a few hundredths of a point.

AFTER-TAX GAIN FROM HARVESTING, LIQUIDATED AT THE END (PP/YR)+0.1%SEL 5%+0.0%SEL 10%+0.0%SEL 20%+0.1%ULT 5%+0.0%ULT 10%+0.0%ULT 20%
read the autopsy

The lot-selection rule paid over a point a year on SELECT, so the next structural idea was the classic one: when a lot is under water, sell it, keep the exposure through a substitute, and use the loss now. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards on the engine’s own trades from the full-history test, 1999 to 2026, with first-in-first-out lots, the test’s tax rates and a pooled loss carry, and the strictest version of the answer: every deferred gain paid at the end of the window. Harvesting at five, ten or twenty percent under the basis improves the after-tax return by +0.09, +0.03 and +0.01 points a year on SELECT and +0.09, +0.04 and +0.01 on ULTRA, and the best of those went negative in the last decade. The reason is in the engine: a rule that cuts exposure when volatility rises sells its losers itself, so the losses a harvester would bank are mostly banked already. The gate asked for a quarter point on both dials; nothing came close. It stays a good idea for a buy-and-hold book and an empty one for this engine.

full-history test, 1999–2026, $25,000, FIFO lots, 35%/20%, deferred tax paid at the end · SELECT baseline 13.73%/yr after tax, harvest 5/10/20%: +0.09 / +0.03 / +0.01% · ULTRA baseline 15.18%, harvest +0.09 / +0.04 / +0.01% · 2017-26 at 5%: SELECT -0.11, ULTRA -0.18 · no wash-sale window modelled (stated)
B-RANKmeasured

Riding the leveraged funds into the close

Late-day rebalancing pressure failed the trading test.
failed on its ownTested as a strategy on its own, after trading costs.

We measured the last hour on days the Nasdaq had moved by mid-afternoon, across thirteen years of minute data. The push is there and grows with the move; a retail trade on it clears costs by about a basis point and was negative for four of those years.

LAST HOUR, SAME DIRECTION AS THE DAY, BY SIZE OF THE MOVE (BP)+3.4 bp0.5%+5.3 bp1.0%+9.2 bp1.5%+12.2 bp2.0%
read the autopsy

Every daily-reset leveraged and inverse fund on the Nasdaq-100 must rebalance at the close, in the same direction as the day, by an amount known by mid-afternoon. We wrote the rules and the pass mark down before the test ran, so nothing could be tuned afterwards on 3,411 sessions of one-minute QQQ bars from IBKR, 2012-12-28 to 2026-09-02, thresholds fixed in advance. The mechanism shows: the last hour continues the day’s direction and the effect rises with the size of the move, from about 3 basis points on half-percent days to about 12 on two-percent days, and a placebo of random sessions never matches it. Then the honest part: the version a cash account can hold, buying at three o’clock on up days and selling at the close, nets 1.3 basis points an event after a dollar and a basis point each way, 0.57% a year, and it lost money from 2013 to 2016. The both-direction version does better (1.97% a year) and needs a short. By the registered gate that is era-dependent, not a product. What the study did settle for us: after a session down one percent or more, the fund is on average 9 basis points above the prior close by 09:45, so the engine’s morning execution of a cut is not paying a late-day penalty.

QQQ one-minute bars, 2012-12-28 to 2026-09-02 · late return same direction as the day: 0.5% +3.4 bp · 1% +5.3 · 1.5% +9.2 · 2% +12.2 (placebo P ≥ 0.996 at every step) · long-only at 1%: 584 events, +1.3 bp net, +0.57%/yr, 2013-16 -1.5 / 2017-20 +5.2 / 2021-26 +0.3 bp · both directions +1.97%/yr · after a ≥1% down day, prior close to 09:45: +8.7 bp
B-RANKmeasured

A diversified ETF book instead of the engine

A diversified ETF book failed the matched-risk test.
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We built a fourteen-fund book with volatility sizing and trend rules and ran it against the engine on one executor. At the same risk it earned less, in every period we split the history into.

SHARPE, 2007–2026, ONE EXECUTOR FOR EVERY ARM, $25,0000.6460/400.51book0.44+sizing0.55+trend0.88engine, same vol
read the autopsy

The second-opinion review of the product asked the right question: is the engine’s edge over a plain diversified book real, and if a broad book beats it, which layer pays, the diversification, the sizing or the timing? Pre-registered before any number was run: a frozen universe of fourteen liquid ETFs from 2007, four arms measured separately. A static book with fixed block weights; the same book scaled to a 10% volatility target; that book with each ETF held only above its own 200-day average; and the engine. The first pass was audited the same day by the Lab, which found four execution defects (a warm-up read as an exit, a fixed cash yield for the engine only, a modelled levered leg, a risk match that did not preserve Sharpe); a second audit then found stale deferred orders, an extra session of delay for the engine and a fee denominator error; this stone is the corrected third pass, every arm on one executor and one timing convention (prior-close signal, fill at that session’s close): decision at the close, fill at the next close, whole shares, a dollar a leg, a cash account that cannot spend the day’s sale proceeds, bills bought and sold like any line, the engine run from its own signal and dial through that same executor with the real levered fund. The static book lost to a two-fund 60/40 (0.51 against 0.64): the real-asset block cost more than it diversified. Sizing to a volatility target made it worse (-0.07). The trend filter helped (+0.12), but by cutting volatility and drawdown, not by picking months: in stocks the months it dropped earned more than the months it kept. The engine at the same volatility beat the best of the four in the full window and in each of the three sub-windows. The precise conclusion, and only this: this frozen construction failed its declared gates under this execution model. It does not show that diversified trend is exhausted, nor that the engine has alpha; it shows that the registered falsifier, a static book matching the engine at the same volatility, did not fire. The window has no dot-com bear, which is stated so this is never read as a full-cycle claim. Status: independently reproduced by the Lab on 2026-09-12 after a five-round audit (stale deferrals, timing, fee accounting, calendar edges, all fixed and tested); the rejection stands under its documented model, with the window’s limits as stated. The one executor is proven by substitution in the archive, and the one deliberate difference between arms is decision frequency: monthly for the books as registered, every session for the engine as sold.

2007–2026 (the ETF histories; no dot-com bear in this window, so not a full-cycle claim), $25,000, one executor for every arm · Sharpe: 60/40 0.64 · static book 0.51 · + vol sizing 0.44 · + trend 0.55 (worst fall -12%) · the engine 0.88, at the same vol 0.88 (worst fall -8%) · sub-windows won by the engine 3 of 3 · kept vs dropped months, stocks +0.82% / +1.02%, bonds +0.40% / +0.28%, real assets +0.83% / +0.66% · at $5,000 every arm reads lower (whole shares)
B-RANKmeasured

Levered funds in the sleeves

a daily-reset 2× bitcoin fund turns a 26%-a-year sleeve into an 11% one; the leverage is eaten by the volatility it is applied to
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Two-times and three-times funds in the sleeves. Daily leverage on an asset that swings sixty percent a year bleeds every day.

THE 20% SLOT ALONE, SINCE 201826.2%-67%btc 1x11.2%-92%btc 2x10.4%-30%gold 1x12.7%-59%gold 2x
read the autopsy

Hold the bitcoin seat in a 2× or 3× daily-reset fund instead of the spot fund, the gold seat in 2× or 3×, gate unchanged. The real 2× funds from their listings, modelled before that from spot with financing and fee; the 3× versions modelled throughout, since no US 3× bitcoin or gold fund exists. Against the shipped pair at equal drawdown, since 2018 and since 2022. Bitcoin levered is off the ladder in both directions: the gated 2× sleeve on its own earns 11% a year against 26% for the 1×, with a −92% drawdown and a −43% worst day, because a daily-reset fund on an asset with 60% volatility pays the volatility decay every day and the gate cannot save it from the days inside a month. The pair with 2× bitcoin has a −37% worst drawdown against −24%, and its extra return is the size dial paid for twice. Gold 2× is a wash since 2018 and half a point since 2022; the modelled 3× gold reads well, and it is a fund that does not exist, on a window that is gold’s bull market. The sleeve fractions are the leverage dial for these seats, and they are already where the frontier put them.

vs shipped 60/20/20 at equal DD, since 2018 / since 2022 · bitcoin 2× in: off the ladder (26.6%/−37.0%) / −0.66 · bitcoin 3× modelled: off the ladder (−49.5%) / off · gold 2× (UGL): −0.24 / +0.65 · gold 3× modelled: +1.41 / +3.64 (no such fund) · sleeve alone since 2018: bitcoin 1× 26.2%/−67%/worst day −21%; bitcoin 2× 11.2%/−92%/−43%; gold 2× 12.7%/−59%/−21%
B-RANKmeasured

Shorting the sleeve while it is out

the inverse fund pays in one bear and bleeds through every recovery the gate is late to
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Betting against a sleeve while it is out, as with grave 110. Same answer: being out is protection, betting on the fall is a new risk.

VS THE PAIR AT EQUAL DD, 2018 / 2022btc -1x out '1-1.21btc -1x out '2+1.57btc -2x out '1-6.36gold -1x out '-2.20gold -1x out '-2.31
read the autopsy

When a sleeve’s gate is off, hold the inverse fund instead of T-bills: −1× or −2× bitcoin, −1× or −2× gold, daily reset, real funds where they exist and modelled before. The NASDAQ version of this idea is grave 88; bitcoin’s bears are three times deeper, so it earned its own run. Against the shipped pair at equal drawdown, since 2018 and since 2022. Bitcoin short-while-out loses 1.2 points since 2018 and wins 1.6 since 2022, which is the 2022 bear and nothing else: the sleeve alone earns 11% a year with the inverse leg against 26% in cash, because the gate is off for months after every bottom while bitcoin doubles, and an inverse fund is on the wrong side of every one of those months. The −2× version loses six points since 2018. Gold short-while-out loses two to four points on both windows; gold’s time below its average is drift, not decline. The gate’s job is to not own the fall. Betting on the fall is a different job, and the same rule is not good at it.

vs shipped at equal DD, since 2018 / since 2022 · bitcoin 1× in, −1× out: −1.21 / +1.57 · −2× out: −6.36 / +2.04 · 2× in, −1× out: off the ladder / +0.90 · gold −1× out: −2.20 / −2.31 · −2× out: −3.61 / −3.84 · both 2× in and −1× out: off the ladder (−38.4%) / +1.77 · sleeve alone since 2018: bitcoin with the inverse leg 11.4%/yr vs 26.2% in cash; gold 5.3% vs 10.4%
B-RANKmeasured

Bitcoin basis carry

the monthly CME basis a regulated account can capture averages 2.7% a year before costs, goes negative four cycles in ten, and pays less than a Treasury bill on the capital it ties up
failed on its ownTested as a strategy on its own, after trading costs.

Earning the bitcoin futures premium. At retail contract sizes the trade does not fit the account.

RETURN ON CAPITAL, %/YRbills+2.00basis @40% mgn+0.462018-9.302022-6.602024+7.90
read the autopsy

The trade every crypto desk describes: buy the spot fund, sell the bitcoin future against it, and collect the premium as the future converges at expiry, with no view on direction. The version a margin account at a US broker can actually run, pre-registered: the front CME contract sold on the first day of each contract month from the exchange’s launch in December 2017, held to its last-Friday expiry, spot-fund fee, slippage, a tracking haircut, and $3 a round trip on the micro contract; capital is the spot plus the 40% margin the broker holds against the short. The entry basis averaged 0.22% a cycle, 2.7% a year, and 42 of 104 cycles were in backwardation, where the trade pays to be in it. Net of costs it earned 0.46% a year on capital against 2% for bills, with 2018 at −9% and 2022 at −7% of capital, the years bitcoin fell. It is uncorrelated with everything, which is true and does not help: beside the pair it loses at equal drawdown at every size and whether the capital comes from the engine or from the margin account. The famous double-digit basis lives in the quarterly contracts and in offshore perpetual funding, in the months bitcoin is running, and neither is on offer to this account. What is on offer is a small premium that disappears in the years you would want it.

CME front month, 104 cycles since the exchange listed bitcoin futures in December 2017, 2018–2026 · mean entry basis 0.22%/cycle (2.7%/yr), 42 negative · net by year: 2018 −13.0%, 2019 −0.5, 2020 +3.8, 2021 +6.0, 2022 −9.2, 2023 +4.4, 2024 +11.0, 2025 +1.9, 2026 +1.2 (on spot notional) · return on capital at 40% margin 0.46%/yr, Sharpe 0.13 · beside the pair at equal DD: 10% from the engine −1.29 / −0.75, 20% −2.46 / −1.64, 20% overlaid −1.24 / −0.79 · one micro unit today ~$11,300
B-RANKmeasured

Gold as an engine

gold’s volatility does not forecast its next month the way the NASDAQ’s does; a vol target on gold reproduces the slice
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Gold as an engine of its own. Gold trends, but too slowly and too rarely to carry a whole book.

VS THE PAIR AT EQUAL DD, 2018 / 2022target .26 '18-0.36target .26 '22+0.72median rv '18+0.00median rv '22-0.94
read the autopsy

The product’s one real relationship, realised volatility predicting next-month risk, applied to the gold seat: instead of a fixed fifth of the book in gold, run a volatility target on it with the 1× and 2× gold funds as legs, gated by gold’s own 200-day, sized like the core, weights decided a day ahead, turnover charged. Two targets, the core’s own 0.26 and gold’s expanding median so the average exposure sits near one. Pre-registered against the shipped pair at equal drawdown, since 2018 and since 2022. The median-target engine lands on the shipped slice to the hundredth since 2018 and loses a point since 2022; the hotter target loses a third of a point since 2018 and earns seven tenths since 2022, which is the dial. The worst day improves, −6% against −10%, and that is the whole effect: a vol target on an asset whose volatility does not predict its return is a smoother way to hold the same thing. The relationship is the NASDAQ’s, and it did not travel.

vs shipped 60/20/20 at equal DD, since 2018 / since 2022 · gold engine, target 0.26: −0.36 / +0.72 · target = median rv: 0.00 / −0.94 · the two sleeve engines together: −0.70 / −2.14 · average exposure 0.93× · sleeve worst day −6.2% vs the slice’s −10.3% · kill: +0.5 on both windows
B-RANKmeasured

Bitcoin as an engine

sizing bitcoin by its volatility holds the least of it in exactly the months it pays, and the 2× fund’s cost eats the rest
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Bitcoin as an engine of its own. Its swings are too large for a volatility dial to tame at any reasonable exposure.

VS THE PAIR AT EQUAL DD, 2018 / 2022median rv '18-0.78median rv '22-2.03target .26 '22-2.37
read the autopsy

The same idea on the bitcoin seat: a volatility target on bitcoin with a spot leg and a 2× leg, the real 2× fund from 2023 and a modelled one before it with its financing and fee, gated by bitcoin’s own 200-day. At the core’s 0.26 target it holds a third of a bitcoin position on average, because bitcoin’s volatility is fifty to sixty, and gives up almost all of the sleeve’s return; at bitcoin’s own median it averages 1.2× and still loses three quarters of a point since 2018 and two since 2022, because bitcoin’s volatility rises into its best months as often as its worst, so the target trims the runs it is meant to ride, and the levered leg’s cost is paid on the way. The worst day improves from −21% to −16%, which is real and not worth the return. Bitcoin’s edge in this product is its drift with the 2022s removed by the gate; the gate does the work, and the sizing should stay fixed.

vs shipped at equal DD, since 2018 / since 2022 · bitcoin engine, target 0.26: off the ladder / −2.37 (18.8%/−19.4% since 2018, a third of the sleeve) · target = median rv: −0.78 / −2.03 · average exposure 1.18× · sleeve worst day −15.8% vs −21.2% · 2× leg real from 2023-06, modelled before
B-RANKmeasured

Rotating the sleeves by momentum

twelve-month relative strength between bitcoin and gold picks the one that just ran and misses the one about to
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
VS THE PAIR AT EQUAL DD, 2018 / 202230/10 '18-5.2730/10 '22-2.2240/0 '22-3.02
read the autopsy

Instead of a fixed 20% and 20%, split the two seats by twelve-month-minus-one relative strength, read monthly and applied a day late: the winner gets thirty and the loser ten, or the winner takes all forty. The gates still apply. Pre-registered, expected to be a dial, and worse than that: it loses five points a year at equal drawdown since 2018 and two since 2022, and takes the worst drawdown from −24% to −30% and −36%, because the two assets alternate on a timescale shorter than the lookback. Bitcoin leads for a year, the rule concentrates in it, bitcoin’s 2022 arrives with the sleeve at thirty or forty; gold leads through the bear, the rule rotates in as gold’s run ends. Cross-sectional momentum between two assets is a coin that remembers the last flip. The fixed split is the answer precisely because it does not.

vs shipped at equal DD, since 2018 / since 2022 · winner 30 / loser 10: −5.27 / −2.22, DD −29.7 · winner 40 / loser 0: off the ladder / −3.02, DD −36.4 · Sharpe 0.97 and 0.81 vs 1.13
B-RANKvalidated

Implied volatility as the sizing input

the option market’s forecast carries a premium and a lag; the engine sizes better on what the market just did than on what the market is charging for
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Sizing on option-implied volatility instead of realised volatility. Implied volatility carries a fear premium that makes the engine too small too often.

EDGE VS REALISED-VOL ENGINE, PP/YRVXN-1.85VXN scaled-2.61blend 50/50-1.00max(VXN,rv)-1.63SELECT VXN-2.54
read the autopsy

The engine sizes on twenty days of realised volatility. The option market publishes its own forecast every day, the VXN, forward-looking by construction. So size on that instead: the VXN as the estimate, the VXN scaled to the realised series’ average so the level is not the dial, a 50/50 blend, and the larger of the two. Pre-registered; 2001 to 2026, both dials, same target, cap, band and exit. Every arm loses at equal drawdown, by one to three points a year, in every era, on both dials. Implied volatility runs about four points above realised on average because it contains the insurance premium, and it moves later than the realised estimate on the way down, so the engine sized on it holds too little through recoveries and is not sized down any sooner into declines, since the two rise together. The scaled version, which strips the level, loses more, which says the shape is worse and not just the mean. VIX as a cap governor, VIX as a veto, VIX as a thermometer and a vol-of-vol dial were already on this wall; this is the last way to put the option market’s volatility number into the engine, and it is the worst.

2001–2026, pre-tax, equal-DD edge vs the realised-vol engine · ULTRA: VXN −1.85, VXN scaled −2.61, blend −1.00, max −1.63 (eras all negative, −0.6 to −4.9) · SELECT: −2.54 / −3.13 / −1.18 / −2.38 · VXN mean 24.9 vs realised 21.1
B-RANKfull battery

Put skew as a crash dial

the skew is a price, not a forecast; it is highest after the fall, and shuffled skew scores the same
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Option skew as a crash warning. It warns often, and the crashes do not come when it warns.

EDGE AT EQUAL DD, PP/YRULTRA x0.75-1.05ULTRA cap 1-1.63SELECT x0.75-1.03SELECT cap 1-0.91
read the autopsy

The price of crash insurance relative to at-the-money insurance, the put skew, is the option market’s own statement about tail risk. From the real NASDAQ chains, daily since 2012: the implied volatility of the put 10% below the market, thirty to forty-five days out, minus the at-the-money put’s, each backed out of the quoted mid, ranked through its own expanding history and used a day late. When it is in its top decile, size down by a quarter or cap leverage at one. Pre-registered, both dials, placebo of the skew history shuffled by year. Every arm loses at equal drawdown, about a point a year, and the placebo is decisive: on ULTRA the shuffled skew beats the real one 98% of the time with the cap and 35% with the scaling, which means the real timing is worse than random timing of the same size. Skew steepens when the market has just fallen and dealers have repriced the tail, which is when the engine has already cut and the recovery is nearest. Nine ideas from dealer positioning, five from the VIX, and now the skew: the option surface describes the last move well and the next one not at all.

2012–2026, a shorter window because the archived option chains begin in 2012, real chains, condition = skew above its expanding 90th percentile (19.7% of days) · ULTRA: ×0.75 −1.05% (placebo beaten 65%), cap 1.0 −1.63 (placebo 2%) · SELECT: −1.03 (21%), −0.91 (8%) · mean skew 7.4 vol points, top decile above 9.0
A-RANKfull battery

A vol-of-vol dial

‘calm market, volatility of volatility exploding’ is a condition that holds six days a year, and sizing down on those days is a coin flip
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A dial on how fast volatility itself is changing. Interesting to look at, worthless to trade.

EDGE AT EQUAL DD, PP/YR (VVIX)ULTRA x0.75-0.08ULTRA cap 1+1.15SELECT x0.75+0.27SELECT cap 1-0.22
read the autopsy

The textbook auxiliary layer: when current volatility is low but the volatility of volatility is spiking, something is cooking, so size down ahead of it. Pre-registered as written, with the condition defined before the run: twenty-day realised volatility below its own expanding median and vol-of-vol above its expanding 90th percentile, measured two ways — the CBOE’s VVIX from 2007 and a realised version, the twenty-day dispersion of the volatility estimate itself, on the full cycle. Action: target scaled by three quarters, or leverage capped at one. Placebo: the vol-of-vol history shuffled by year. The condition is true on 6% of days with VVIX and 0.4% with the realised measure. One arm of eight reads a point: ULTRA capped at one on VVIX days, +1.15 at equal drawdown, but it loses a fifth of a point on SELECT, the shuffled series beats it 12% of the time, and its three quarters live in 2017–2026. Every other arm is within a quarter point of zero either way. The realised version, the one with twenty-seven years, is inert. The volatility target already reads the volatility of volatility, because a rising estimate is a rising estimate; a dial on the second derivative fires too rarely to matter and, when it fires, does not know the direction.

VVIX 2007–2026, condition 6.4% of days · ULTRA ×0.75 −0.08%, cap 1.0 +1.15 (placebo 88%, SELECT −0.22) · SELECT ×0.75 +0.27, cap 1.0 −0.22 · realised vol-of-vol, full cycle, condition 0.4% of days: −0.02 / +0.10 / −0.07 / −0.01 · kill: both dials ≥ +0.3 and placebo 90%
A-RANKvalidated

Dampening the cut above the trend line

the engine already has an asymmetric band and a trend exit; letting the cut wait while the trend is up adds drawdown and a rounding of return
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Letting cuts wait while the trend is still up. The fall it lets through costs more than the trades it saves.

EDGE AT EQUAL DD, PP/YRULTRA cut 10%+0.08ULTRA cut 15%+0.27SELECT cut 10%-0.71SELECT cut 15%+0.13
read the autopsy

The other textbook layer: markets fall faster than they rise, so when volatility spikes while price is still above its 200-day average, dampen the selling to avoid whipsaws. The engine already has the two ingredients — a band that adds at 25 to 50% and cuts at 5 to 7.5%, and the monthly 200-day exit — so the untested piece was widening the cut to 10% or 15% while above the line. Pre-registered, full cycle, both dials. The wider cut earns a tenth to a quarter of a point at equal drawdown, below the bar, with the eras disagreeing: it loses in 1999–2008 and 2009–2016 and wins in 2017–2026, where volatility spikes above the trend line were the ones that resolved upward. The cut fires when volatility has already risen, which is information; waiting because the price is still above a moving average trades that information for a bet that the spike is noise, and over twenty-seven years the spikes above the line were not noise often enough.

full cycle, pre-tax, equal-DD edge · ULTRA (cut 7.5%): cut 10% above the 200d +0.08, cut 15% +0.27 (eras −0.14 / −0.81 / +1.23) · SELECT (cut 5%): 10% −0.71, 15% +0.13 (eras +0.32 / +0.47 / +1.48, DD −22.6 vs −21.3) · kill: both dials ≥ +0.3
B-RANKvalidated

A cleaner volatility input

the estimators that use the day’s range are quieter and wrong about the thing that matters, which is the overnight gap
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A cleaner, more careful volatility measure. It changes the number and not the decision.

EDGE AT EQUAL DD, PP/YR (ULTRA)Parkinson-1.55Garman-Klass-1.57Yang-Zhang-2.40
read the autopsy

The engine sizes on twenty days of close-to-close volatility. Range-based estimators — Parkinson, Garman-Klass, and Yang-Zhang, which adds the overnight gap back — use each day’s high and low and are several times more efficient in the textbook sense. Pre-registered on the full cycle, both dials, same target, same cap, same band. Parkinson and Garman-Klass read the NASDAQ’s volatility at 19.4% against 22.8% close-to-close, because the range inside the session does not contain the gap and the gap is most of the NASDAQ’s variance; the engine then holds more and the extra return is the dial, minus 1.5 to 2.6 points at equal drawdown. Yang-Zhang matches the close-to-close level and still loses 2.4 points at equal drawdown, negative in the two later eras, because its weighting of the overnight and intraday pieces changes the estimate exactly on the days after a gap, which are the days the sizing has to be right. The noisy estimator is the one that measures the risk the engine actually carries.

full cycle, pre-tax, equal-DD edge · ULTRA: Parkinson −1.55, Garman-Klass −1.57, Yang-Zhang −2.40 (eras +1.06 / −2.82 / −1.33) · SELECT: −1.46 / −2.56 / −2.49 · estimator means 22.8 / 19.4 / 19.4 / 23.1 · trade counts unchanged
B-RANKvalidated

Tax-aware cuts

waiting for a deeper cut to avoid a short-term gain costs more in the market than it saves at the IRS
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Delaying cuts to avoid a tax bill. The market does not wait for the tax year.

AFTER-TAX EDGE VS PLAIN CUT, PP/YRULTRA +2.5pt-0.29ULTRA +5pt-0.03SELECT +2.5pt-0.08SELECT +5pt+0.19
read the autopsy

With lots sold highest-cost-first, the setting the product ships, a cut sometimes sells shares bought within the year, and the gain is taxed at the higher rate. So when the shares a cut would sell are mostly young, wait: require the cut to be two and a half or five points deeper before acting. Pre-registered, full cycle, after tax with the 35/20 rates and loss carry. It loses on ULTRA at both settings and on SELECT at the shallower one; the deeper wait scrapes a fifth of a point on SELECT at a point and a half more drawdown, below the quarter-point bar. The cut exists to sell before a decline runs, and a cut that waits is a cut that is late; the tax saved on a few young lots is smaller than what the delayed sale gives back. The tax-lot setting itself is confirmed on this run at +0.6 and +1.3 points a year, which is the part of this idea that was ever real.

after tax, full cycle, vs the same engine with the plain cut · ULTRA: +2.5pt wait −0.29%, +5pt −0.03 · SELECT: −0.08, +0.19 (DD −1.4) · the tax-lot setting alone +0.60 / +1.26 · kill threshold 0.25
B-RANKmeasured

The gold sleeve on real yields

real yields explain gold; they do not time it better than gold does
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Holding the gold sleeve only when real interest rates say to. The gold rule already knows; the extra rule mostly agrees with it.

GOLD SLEEVE EDGE VS ULTRA, 2005-26price gate+2.21real-yield gat+1.38both+1.23either+2.08no gate+2.61
read the autopsy

Gold’s textbook driver is the real interest rate, and its twenty-year bear ran alongside high ones. So gate the sleeve on the ten-year TIPS yield instead of gold’s own price: hold gold while the real yield is below its own 200-day average, read monthly, or require both signals, or accept either. The Treasury’s daily real-yield curve from 2003, the same one-day lag as every other signal, beside ULTRA at a fifth of the book. The real-yield gate earns a point and a half at equal drawdown against the price gate’s 2.2, requiring both is worse still, and accepting either merely reproduces the price gate. The two gates agree on 44% of days, and where they disagree the price knows first: in the 2005–2008 stretch gold rose while real yields were flat and the yield gate sat out most of it. Holding gold with no gate at all beats every gate on this window, which is the cost the price gate pays on purpose for the 1980–2001 bear that this window does not contain. Real yields are a good explanation and a slow signal.

ULTRA 80 + gold 20, 2005–2026, equal-DD edge vs ULTRA · gold above its 200d (shipped) +2.21 · real yield below its 200d +1.38 · both +1.23 · either +2.08 · no gate +2.61 · eras (shipped / real yield): +1.83 / +0.36, +1.68 / +2.47, +1.95 / +2.13 · gates agree 44% of days, corr +0.31
B-RANKfull battery

Futures positioning as a dial

when speculators are most long the NASDAQ, the NASDAQ has mostly kept going up; cutting exposure there is paid for with the best months
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Using futures positioning reports as a dial. They tell you what happened, not what is next.

EDGE AT EQUAL DD, PP/YRULTRA cap>p90-1.68SELECT cap>p90-0.77ULTRA x.75>p90-0.89SELECT x.75>p9-1.14mirror <p10+0.05
read the autopsy

Twenty-six years of the CFTC’s weekly Commitments of Traders report for NASDAQ 100 futures, the non-commercial net position as a share of open interest, ranked through its own expanding history so nothing is known in advance, and used from the Monday after the Friday release. Three pre-registered arms: cap leverage at one when speculators are in their top decile, scale the target down by a quarter there, and the mirror, cap when they are in the bottom decile. Every crowd-is-long arm loses at equal drawdown, by three quarters of a point to nearly two points a year on both dials, and the mirror does nothing. The placebo settles it: shuffle the positioning history by year and the real series does worse than 82% of the shuffles, so the rule is not merely useless, it is timed against the market. Speculators are long the NASDAQ when it is trending, and the engine already knows when it is trending. Crowd positioning is a description of the last quarter’s trend wearing the costume of a contrarian signal.

COT since 2000, 2000–2026, 1,368 weeks, expanding percentile · cap 1.0 above the 90th percentile: ULTRA −1.68% at equal DD, SELECT −0.77 · target ×0.75 above the 90th: −0.89 / −1.14 · mirror below the 10th: +0.05 / −0.11 · placebo (years shuffled, 100 draws): real beats 18%
B-RANKvalidated

Tax-loss harvesting with a twin fund

the engine’s losses are already deductible; a twin only rescues the few the wash-sale rule would have deferred, and that is a few hundredths of a point
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Selling a losing fund and buying its twin to bank the tax loss. The engine's own cuts already bank most of the losses.

AFTER-TAX, PP/YRwash rule cost-0.35wash rule cost-0.32twin gains U+0.02twin gains S+0.08
read the autopsy

Pre-registered and run on the full 1999–2026 cycle with lots, holding periods, the 35/20 rates and loss carry: first the wash-sale rule itself, which the published after-tax numbers had never modelled — a loss is disallowed when the same fund is bought within thirty days either side and goes into the replacement lot’s basis — and then the fix, buying a twin fund with identical returns for thirty days after any loss sale so the loss stands. The wash rule costs the engine about a third of a point a year on both dials, which is how optimistic the after-tax figures on this site were, and it is now stated there. The twin recovers almost none of it: two hundredths of a point on ULTRA, eight on SELECT, never more than a fifth of a point in any era. The reason is that the engine’s loss sales are the exit and the cuts, and the buys that would wash them come weeks or months later, outside the window; the deferred losses are a rounding of the tax bill, not a lever on it. The tax-lot setting shipped in September is worth a point; this was worth checking and is worth nothing.

after-tax CAGR, full cycle · ULTRA: published model 12.49 → with the wash rule 12.14 → with the twin 12.17 (+0.02) · SELECT: 10.47 → 10.15 → 10.23 (+0.08) · eras +0.04 / +0.07 / +0.01 and +0.09 / +0.21 / +0.01 · days holding the twin 18% / 27% · kill threshold 0.25%
A-RANKfull battery

Post-exit volatility shrinkage

by the time the exit lets the engine back in, the volatility estimate has already normalised; the rule never binds
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Shrinking the volatility estimate right after an exit. By the time the rule lets the engine back in, the estimate has already settled.

EDGE AT EQUAL DD, PP/YRULTRA M=1..3-0.03SELECT M=1..3+0.19
read the autopsy

The idea: after a crash the twenty-day volatility estimate stays swollen for weeks, so the engine holds its smallest position exactly when recoveries are steepest, April 2020 being the poster child. So for one, two or three months after each re-entry, cap the estimate at its long-run median. Pre-registered, full cycle, both dials, with the placebo of the same cap in random months. The shipped exit re-enters three times in twenty-seven years — May 2003, June 2009, April 2023 — and each time it re-enters after two confirming monthly closes above the band, by which point realised volatility is already at or below its median. The cap binds on a handful of days and changes ULTRA by three hundredths of a point; on SELECT it earns a fifth of a point, all of it from the 2009 re-entry, which the placebo beats anyway. April 2020 is not on the list at all, because the confirm never let the engine leave that spring. The premise was true of a fast exit; the shipped exit is slow on purpose, and the slowness already does this job.

re-entries 2003-05, 2009-06, 2023-04 · ULTRA M=1/2/3: 17.06/−27.4 vs plain 17.01/−27.2, equal-DD −0.03%, per-re-entry gain 0.0% each · SELECT +0.19%, 100% from 2009, placebo beaten 93–97% on a fifth of a point · eras 0.00 / +0.01 / +0.05
B-RANKmeasured

The engine on the semiconductor root

the highest-volatility broad index is also the one whose crashes are deepest and longest; the rule cannot size around a dot-com bust that never recovers
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The engine on semiconductors instead of the Nasdaq-100. More return, far deeper falls, and a worse trade of one for the other.

2001-2026, ULTRA DIAL13.1%-40%semis17.7%-27%NASDAQ
read the autopsy

The same law, the same exit and the same dials on the semiconductor index instead of the NASDAQ 100: the sector fund as the base from 2001, the 3× fund as the levered leg from its 2010 listing and modelled before that from the index with financing and fee. Pre-registered against ULTRA on the NASDAQ at equal drawdown over the same window. It loses by eight points a year on the hot dial and near seven on SELECT at equal drawdown, its worst drawdown is −40% against −27%, and the first era is negative outright: the semiconductor index fell 80% from 2000 and did not see its high again until 2017, so an exit that sells after two bad monthly closes and re-enters after one good one spends that decade paying for whipsaws in a market that was going nowhere. It wins the last era, 2017 to 2026, by two and a half points at a deeper drawdown, which is the sector’s bull market and nothing the engine did. The futures work showed the engine’s edge is a high-volatility-root edge; this is the reminder that the root also has to trend for a quarter century.

2001–2026, pre-tax · ULTRA dial: SEMI 13.1%/−40.4% vs NASDAQ 17.7%/−27.2%, −8.0% at equal DD; eras −2.1 vs +7.6, 12.7 vs 21.0, 27.1 vs 24.5 · SELECT dial: 11.5/−31.9 vs 15.1/−21.3, −6.7% · exits 0.35/yr · 3× leg modelled 2001–2010
B-RANKvalidated

A drawdown throttle

cutting exposure after a drawdown is the dial again, applied at the worst possible moment
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
FULL CYCLE, PRE-TAX13.8%-23%throttle U17.0%-27%plain U12.7%-19%throttle S14.5%-21%plain S
read the autopsy

Scale the target down by a quarter when the account is more than 10% below its peak and by half below 20%, restoring it at a new high. Pre-registered as the idea most likely to be a dial, and it is: at equal worst drawdown it loses 1.2 points a year on ULTRA and 0.7 on SELECT, and in the two later eras it loses seven to eleven points raw, because a drawdown from peak is where the engine’s recoveries begin and the throttle holds half a position through all of them. The engine already has a drawdown response: the volatility target, which cuts exposure when the market gets rough and restores it when it calms, on the market’s own information rather than the account’s. A rule keyed to the account’s own high-water mark knows nothing the market does not and forgets it later.

full cycle, pre-tax · ULTRA 13.8/−23.1 vs plain 17.0/−27.2, −1.15% at equal DD, eras −1.7 / −10.2 / −10.7 raw · SELECT 12.7/−19.3 vs 14.5/−21.3, −0.73%, eras −2.2 / −7.7 / −9.4
B-RANKmeasured

Delta-neutral variance carry

after the bid and the hedging, the volatility premium on the NASDAQ is smaller than a Treasury bill, and it is paid back in the crash months
failed on its ownTested as a strategy on its own, after trading costs.

Selling option variance while hedged. The premium is real and the tail eats it.

CARRY ALONE ON CASH, %/YRbills+2.000.25x notional+1.780.5x+1.601.0x+1.18
read the autopsy

Sell the at-the-money NASDAQ straddle every month, thirty to forty-five days out, both legs at the bid, and hedge the direction away every day with shares using the delta from the implied volatility in the quotes, so the return comes only from implied volatility exceeding realised volatility, not from where the market goes. Fourteen years of real chains, daily marks, expiry settled at intrinsic, hedge commissions and half a cent a share. The premium is there — implied averaged 19.4 against 18.5 realised — and it is too thin to survive its own costs: run on a book of cash, the strategy earns 1.2 to 1.8 percent a year against 2.0 for the bills it sits on, at drawdowns of 3 to 15 percent depending on size, with two hundred and sixty hedge trades a year. The good years pay a point or two; 2018, 2020 and 2022 each take it back, because a short straddle is short gamma and the realised volatility of a crash month exceeds anything implied the month before. Beside the engine it is worse than useless: it lowers return, deepens the worst drawdown by one to three points, and makes the worst days worse, since the months it loses are the months the engine loses. As for the account: one NASDAQ straddle is $72,000 of notional and about $14,000 of margin, so a $25,000 account cannot hold one. A real premium, priced fairly, hedged honestly, and smaller than the cost of collecting it.

2012–2026, real chains, daily delta hedge · carry alone on cash: 0.25× notional 1.78%/yr, DD −3.2%; 0.5× 1.60%, −6.6%; 1.0× 1.18%, −15.0% · bills 2.0% · mean IV 19.4 vs RV 18.5 · worst months 2020-03 −10%, 2025-04 −9%, 2022-11 −6% at 1.0× · beside ULTRA at equal DD: −1.2% (0.25×), −2.7% (0.5×); worst-5% days −288 → −306 / −320bp · monthly corr with the NASDAQ +0.34 · one straddle = $72k notional, ~$14k margin
B-RANKfull battery

A convex overlay: buying put spreads so the engine can run hotter

the premium costs two and a half points a year and buys a third of a point of drawdown; the engine already has a cheaper exit
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Buying put spreads so the engine could run hotter. The insurance costs more than the extra exposure earns.

2012-2026, PRE-TAX20.7%-26%overlay U23.2%-27%plain U16.3%-21%overlay S19.4%-21%plain S
read the autopsy

Pre-registered before the run, then run exactly as written: every month buy a NASDAQ put spread, long 5% below the market and short 15% below, thirty to forty-five days out, spending half a percent of the account in premium, paying the ask and receiving the bid, marked daily at the mid on fourteen years of real chains, on SELECT and on ULTRA with their shipped exits. The idea was not income; it was insurance that would let the engine run a hotter dial at the same worst drawdown, with the restored return paying for the premium. It does not. The spread costs two and a half to three points a year and shallows the worst drawdown by a third of a point, so raising the engine’s dial to match buys back almost nothing: minus 1.8 points a year on ULTRA at equal drawdown, and three points raw on SELECT. The mechanism is real and useless at once: shuffle the monthly payoffs across months and the real timing is shallower than 94% of the draws — the spread does pay in the crash months — but 2022 is the only year it paid at all, and it took every other year’s premium to pay for it. The engine’s monthly exit already leaves the market in those same months, for free. Insurance bought on top of an exit that already works is a second premium for the same event. Whole contracts change nothing at $25,000 or $100,000; the structure was simply too expensive at every size.

2012–2026 because the archived chains begin in 2012, pre-tax, 0.5% of NLV a month · ULTRA 23.2%/−26.9% → 20.7%/−25.9% (raw −2.5%, equal-DD −1.8) · SELECT 19.4% → 16.3% (−3.1 raw) · premium ~28% of starting NLV a year on ULTRA, payoffs half of it back; net positive in one year of fourteen (2022) · random-day placebo: real beats 79% on CAGR, 85% on Sharpe, 49% on DD · 2022 the only paying year · whole contracts at $25k / $100k within 0.2% of fractional
B-RANKmeasured

Dividend and value equity ballast

a dividend fund is the stock market with a value tilt; the engine already owns the stock market, and the yield is not extra return
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Dividend and value funds as ballast. They carry the same stock-market risk the engine already holds, and pay less for it.

20% SLEEVE, EDGE VS ULTRA (OWN WINDOW)SCHD-0.40VYM-0.80HDV+0.20NOBL-1.40QYLD-1.10
read the autopsy

Seventeen dividend and value equity funds with at least eight years of history — dividend quality, high dividend, aristocrats, dividend growth, low-volatility dividend, international dividend, and the older covered-call index funds — each with distributions reinvested, put through the construction the bitcoin-and-gold pair came through: a fifth of the book in the fund, either held only while it is above its own 200-day average read monthly or held always as ballast, beside the engine dialled to the same drawdown, on each fund’s own window from its listing (2015 for most) and since 2022, then added at 5% and 10% beside the pair itself. Every one lands between minus two points and plus half a point a year at equal drawdown beside the engine. Added beside the pair, the best of them earns a third of a point on its own window and nothing since 2022, and every one deepens the worst days against the same money in bills, because a dividend fund falls with the market on the days that matter: held always, its worst-day correlation is three quarters, the same as the index; gated, it is a tenth, because it is in cash. The yield is paid out of the price, so a fund that yields 4% and returns 11% is an 11% fund. Beside an engine that already holds the NASDAQ at up to twice its size, a slower piece of the same market is dilution with a coupon. Not covered by this stone, because they are not the same mechanism and do not have the history: the premium-income funds JEPI (from 2020) and JEPQ (from 2022) were run on their own short windows only as a diagnostic — JEPI reads minus a point beside the engine, JEPQ plus half a point at a correlation of 0.96, which is the engine with its upside sold — and neither is validated either way. The single-stock option wrappers and the 2025 listings (the YieldMax funds, AMDW, CHPY, BLOX) are out of scope as ballast: leveraged or option-wrapped exposure to one stock, a semiconductor basket, or crypto, one to three years old, each of which trails the thing it wraps by ten to twenty points a year at a correlation above 0.95. They are a bet on the underlying with the upside sold, and there is no eight-year test they could have taken.

20% sleeve, equal-DD edge vs ULTRA, own window / since 2022, gated | always · SCHD −0.4 / −1.1 | +0.5 / +1.1 · VYM −0.8 / −0.3 | 0.0 / +1.3 · HDV +0.2 / +1.3 | −0.2 / +2.4 · NOBL −1.4 / −1.6 · QYLD −1.1 / +0.5 · added at 10% beside BTC + AU: best +0.4 / −0.1 (FDVV), most negative · worst-day cost vs bills −7 to −27bp · gold on the same run: +1.8 / +2.9 gated, +3.1 / +4.0 always · diagnostics only: JEPI −0.9 (2021–26), JEPQ +0.5 (2023–26, corr 0.96); NVDY 58.5%/yr vs NVDA 86.4%, TSLY 8.1% vs TSLA 19.3%, CONY 10.1% vs COIN 29.7%, MSTY 15.0% vs MSTR 24.1%, AMDY 43.2% vs AMD 66.4%, all corr 0.95–0.98
B-RANKmeasured

Sizing the sleeves by their own volatility

it is the sleeve-size dial wearing a volatility badge
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Sizing each sleeve by its own volatility instead of a fixed slice. The dial adds trades and takes away return.

PAIR EDGE AT EQUAL DD, SINCE 2018fixed 20/20+3.20vol tgt .30+2.70vol tgt .45+2.70vol tgt .60+2.70vol tgt .80+3.20
read the autopsy

The engine sizes its core by trailing volatility, so give the bitcoin and gold sleeves the same treatment: each slot scaled down when its own twenty-day volatility runs above a target, capped at the full slot, four targets from 30% to 80%, on the lagged full-history test since 2018 and since 2022, against the engine dialled to the same drawdown. Every target lands on or just below the fixed pair’s own frontier: a tighter target holds less bitcoin and earns less at a smaller drawdown, a looser one converges on the fixed 20/20, and none of them beats the fixed sleeves at equal drawdown. Volatility sizing pays inside the core because the NASDAQ’s volatility forecasts its next month; a fifth of a book in bitcoin is already a small, blunt position, and shrinking it further in its noisy months buys nothing the trend rule has not bought. A mechanism that reproduces a dial is a dial.

pair 60/20/20 fixed, since 2018: 22.2%/yr −24% DD, +3.2% at equal DD; since 2022 +6.1 · vol target 0.30 / 0.45 / 0.60 / 0.80: +2.7 / +2.7 / +2.7 / +3.2 since 2018, +3.9 / +5.1 / +6.1 / +6.1 since 2022 · none above the fixed pair
B-RANKmeasured

A trend basket beside the engine

the diversifiers dilute the two streams that pay, and there is no third stream in the rest
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A basket of trend-following funds beside the engine. It ties cash on return and loses on the worst days.

EDGE AT EQUAL DD, SINCE 2018btc+gold pair+3.20basket 1/N+0.20basket 1/n-on+1.00inverse-vol+0.00rest, no btc/a-0.40
read the autopsy

A home-made managed-futures sleeve: one slice of the book spread across whichever non-equity assets sit above their own 200-day average read monthly — gold, bitcoin, broad commodities, 7-to-10-year and 20-year Treasuries, silver, copper, platinum — weighted equally, or equally among the ones that are on, or by inverse volatility, at a fifth and at two-fifths of the book, against the engine dialled to the same drawdown, since 2018 and since 2022. Every basket earns less than the plain bitcoin-and-gold pair at equal drawdown, the best by two points a year since 2018, because six of the eight assets go nowhere on their own and every dollar spread onto them is a dollar taken from the two that do. Run the basket without bitcoin and gold, to ask whether a third stream is hiding in the rest, and it reads between minus half a point and plus half a point. Carve a slice for it out of the engine beside the pair and it costs one and a half to three points. Diversification across trend streams is what a futures book does with leverage and forty markets; done with eight spot funds and no leverage it is dilution. Two streams was the answer; the weekly screen keeps checking whether a third one appears.

pair 20/20: +3.2% since 2018, +6.1 since 2022 · basket 1/N f=.20: +0.2 / +1.8 · 1/n-on f=.20: +1.0 / +1.7; f=.40: +0.3 / +3.2 · inverse-vol f=.40: −0.0 / +0.8 · rest without bitcoin and gold: −0.4 to +0.6 / −0.1 to +1.0 · pair + rest 10%: +1.6 / +5.6 · the screen: 44 funds, 2 marginal survivors (DBC +0.9, IEF +0.6)
B-RANKmeasured

A currency sleeve beside the engine

a currency fund goes nowhere at a pace a fifth of a book cannot feel
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A currency sleeve. Currencies do not trend reliably enough to pay for the seat.

20% SLEEVE, EDGE VS ULTRA 2015-26UUP+0.60FXE+0.10FXY-0.10FXC-1.10gold+1.70
read the autopsy

The bitcoin candidate’s construction with a currency in the seat: a fixed 20% of the book in a currency fund — the dollar long and short, the euro, yen, sterling, the Australian and Canadian dollars, the Swiss franc, a basket of emerging currencies — held only while it is above its 200-day average read monthly, cash otherwise, against the engine dialled to the same drawdown. Every one lands within a point of nothing, from −1.1 to +0.6 a year, since 2015 and since 2022 alike. Held outright the funds return between −3 and +2.5 percent a year with volatility a fifth of bitcoin’s, so a fifth of the book in one of them moves the book by a rounding error whichever way the trend rule leans, and the rule itself has nothing to catch in a series that mean-reverts around zero. Currencies earn their keep in a futures book with leverage and carry across a dozen pairs at once; that book was tested and buried as grave 102. As a spot sleeve they are neither a return stream nor a hedge. A seat needs a different asset that goes somewhere on its own.

20% sleeve, equal-DD edge vs ULTRA, 2015–2026 · UUP +0.6 · FXB +0.2 · FXE +0.1 · FXY −0.1 · UDN −0.3 · CEW −0.3 · FXF −0.9 · FXA −1.0 · FXC −1.1 · gold +1.7 · bitcoin +11.5 · funds held: −2.8 to +2.5%/yr
B-RANKmeasured

An oil sleeve beside the engine

crude pays the roll and trends badly; the energy stocks that do trend owe it to one year
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

An oil sleeve beside the engine. Oil falls with the market and does not rise with it; it added risk and no return.

20% SLEEVE, EDGE VS ULTRA 2015-26USO-0.90BNO-0.20XLE+0.40XOP+2.80gold+1.70
read the autopsy

The bitcoin candidate’s construction with oil in the seat: a fixed 20% of the book in a crude or energy fund, held only while it is above its 200-day average read monthly, cash otherwise, executed the day after the read, against the engine dialled to the same drawdown. The crude funds — WTI, Brent, the optimised-roll version — earn between minus one point and nothing over their own history since 2015; the trend rule cannot rescue an asset that loses 5 to 10 points a year to the futures roll and spends most of a decade below its average. The energy stocks do better, and one, the producers fund, reads +2.8 points a year; but almost all of it is 2022, the inflation year, and its correlation with the NASDAQ is twice bitcoin’s. Gold, the sleeve the product already has, earns more than every oil variant but one at a third of the correlation. A sleeve has to be a different asset that goes somewhere on its own. Oil is a different asset that goes in circles.

20% sleeve, equal-DD edge vs ULTRA, 2015–2026 · USO −0.9% · BNO −0.2 · DBO +0.1 · XLE +0.4 · XOP +2.8 (2022–2026 +4.3, the only era) · gold +1.7 · bitcoin +11.5 · crude funds held: 0.5–10%/yr with −62 to −87% drawdowns
B-RANKvalidated

A valuation dial from the Shiller P/E

the market has been ‘expensive’ since 1996, so every valuation rule is a thirty-year instruction to own less
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Owning less when the market looks expensive. By that measure the market has looked expensive since 1996, so the rule owned less for thirty years.

S&P SLEEVE, EDGE AT EQUAL DDout top decile-7.40no adds-4.90cap top quinti-1.20
read the autopsy

Use the Shiller cyclically adjusted P/E, through its own expanding percentile so nothing is known in advance, to govern the engine: step out when it is in its top decile, scale the volatility target down as it rises, refuse to add exposure while it is in the top decile, or cap leverage at one while it is in the top quintile. On the S&P fund pair from 1994, its own history, and on STEADY from 2001. Every arm earns less than the engine it governs, most of them by three to seven points a year, because the percentile has sat above 0.9 for most of the last thirty years and the rules simply hold less for decades. The one arm that improves Sharpe, scaling the target down with valuation, is exposed by the placebo: shuffle the P/E history by year and the shuffled version scores the same, so its whole effect is owning less, not knowing when. Valuation tells you the decade’s return. The engine trades the month, and the month does not care.

S&P sleeve: out in top decile −7.4% at equal DD, no-adds −4.9, cap at top quintile −1.2 · vol-scaled 10.1%→7.4%; CAPE shuffled by year gives the same Sharpe (beats 52% of draws) · STEADY: vol-scaled 11.7%→8.5% at −10% DD, +0.5% at equal DD, all three eras lower · CAPE percentile today 0.95
B-RANKvalidated

Parabolic SAR as the exit

it exits twelve times a year to avoid what the monthly rule avoids three times in a quarter century
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A popular chart indicator as the exit. It exits five to sixteen times a year and each exit costs more than it saves.

SELECT, EDGE AT EQUAL DDPSAR .02/.20-14.10slow .01/.10-11.70slower-11.70month-end-13.90
read the autopsy

Replace the engine’s exit — two monthly closes 5% below the 200-day average to leave, one 5% above to return — with Wilder’s parabolic stop, the accelerating trailing stop that flips whenever price crosses it. Standard settings, slow settings, slower still, fast settings, and the same stop read only at month end, all on the full 1999–2026 cycle, on SELECT and ULTRA. Every version loses eleven to fourteen points a year against the engine dialled to the same drawdown, and the fast version loses so much the control cannot reach its drawdown. The reason is the exit count: the parabolic stop leaves the market five to sixteen times a year, the monthly rule about once every nine years, and each false exit is a round trip through the band paid in missed days. A trailing stop is a fine instrument for a single trade with a defined risk. A strategy that is supposed to be in the market for decades cannot afford an exit that is triggered by every ordinary pullback.

SELECT, shipped exit 14.5%/yr −21% DD, 0.1 exits/yr · PSAR 0.02/0.20 daily 4.5%, 12 exits/yr, −14.1% at equal DD · slow 0.01/0.10 6.8%, −11.7% · slower 0.005/0.05 7.5%, −11.7% · fast 0.03/0.30 2.3%, DD −54% · month-end 0.02 5.3%, −13.9% · the dot-com-and-2008 era negative or near zero on every version
A-RANKvalidated

Shorting the market while the engine is out

the exit knows when to leave; it does not know the market will keep falling
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

When the engine leaves, why not bet on the fall? Because leaving is a decision and betting on the fall is a prediction, and the prediction loses.

SELECT, EDGE AT EQUAL DD-1x half, out-1.70vol-gated -1x+1.00gate at 45%-2.20
read the autopsy

The engine already has an oracle for bad markets: the monthly exit that takes it to cash. So when it is out, hold an inverse fund instead of cash — minus one times the index, minus two, minus three — sized as a half or a third of the account, with the daily-reset decay and the fee of the real products modelled, on the full 1999–2026 cycle, on SELECT and on ULTRA. Shorting whenever the engine is out loses at equal drawdown on SELECT and on ULTRA and, above one times, blows the worst drawdown out from the low twenties to the high thirties, because the engine is out for a long time in years that turn out fine, and an inverse fund bleeds every one of those days. Add a second condition, short only when volatility is also high, and it scrapes a point past the control on SELECT and on ULTRA — entirely from 2001 and 2008, negative through 2009–2016, and gone if the volatility threshold moves from 30% to 45%. A gain that lives in two crashes and dies when its one dial is nudged is a fit, not a finding. Random out-days beat nothing here; the real timing is the exit’s, and the engine already owns it.

SELECT: −1x half whenever out −1.7% at equal DD; −1x whole / −2x / −3x DD −38% · ULTRA: −1.0 to −3.4% · volatility-gated −1x: +1.0 / +1.1%, eras +4.4 / −1.2 / +0.4, gate at 45% −2.2 / +0.5 · random out-days Sharpe 0.56 vs real 0.84 (the exit’s value, already owned)
A-RANKmeasured

A bitcoin sleeve on the engines

its whole edge is one era that has not repeated
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A bitcoin sleeve beside the engine. Its edge was 2015 to 2017; since 2018 it adds nothing that the gold sleeve does not add more cheaply.

20% SLEEVE EDGE, 2015-26 VS 2018-26ULTRA '15-26+4.20ULTRA '18-26+0.90SELECT '15-26+2.70SELECT '18-26-0.20STEADY '18-26-0.70
read the autopsy

Give ULTRA, SELECT or STEADY a fifth of the account in a bitcoin fund, or a bitcoin-and-ether fund, run the way the gold sleeve is run: sized to volatility, never levered, stepped aside by the same monthly exit. Bitcoin priced as a spot fund less a 0.25% fee, because the real funds are two years old; ether joins when it lists in late 2017. On the whole window, a shorter window because bitcoin has no price before 2014, the bitcoin sleeve looks like a gift: every engine earns two to four points more a year at the same worst drawdown, and random-sign bitcoin beats it in only 2% of draws. Then split the years. 2015 to 2017, when bitcoin compounded at 85% a year, supplies all of it; 2018 to 2021 is flat to slightly negative; 2022 to 2026 is negative. Rerun from 2018 alone and the sleeve adds nothing at equal drawdown on ULTRA, costs a quarter point on SELECT and most of a point on STEADY. The ether version is worse in every window. A sleeve is only insurance if its good years are the engine’s bad years; bitcoin’s good year was everyone’s good year, and it has not had another. The halving cycle was checked too: the eighteen months after each halving do carry the sleeve’s return, and that phase ranks second or third among the twenty-four offsets the same rule could have used — but the gain in that window has shrunk from 22× to 7.5× to 1.7× across the three cycles the data holds, the current window closed in October 2025, and a rule that says ‘own bitcoin from April 2028’ is a forecast, not a strategy.

2015–2026: ULTRA +4.2%, SELECT +2.7%, STEADY +2.3% at equal DD with a 20% sleeve · eras: all of it in 2015–17 · 2018–2026 alone: ULTRA +0.9 / +0.0%, SELECT −0.2, STEADY −0.6 to −0.8 · bitcoin+ether sleeve negative on SELECT and STEADY in every window · post-halving 18 months: +4.2% at equal DD on ULTRA, the other 30 months −1.8; per-cycle gain 22× / 7.5× / 1.7×
B-RANKmeasured

Switching vehicles by volatility in the futures engine

the calm-market edge of the fund is real, and switching cannot catch it
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The futures engine’s one weak era is a calm, steadily rising market, where a daily-reset 2× fund compounds better than a constant-notional contract. So hold the fund when realised volatility is low and the contract when it is high, at five thresholds, on the full 1999–2026 cycle after tax, at three account sizes. Every threshold loses to simply holding the contract all the time, by half a point to nearly a point a year. Each switch sells fund lots, which is a taxable event, and buys or sells a contract, and the two costs together are larger than the calm-era gap they were meant to close. The lesson is the same one the rest of this wall keeps teaching: a real pattern is not a tradable one when the act of trading it costs more than it pays.

fund below rv 0.15 / 0.20 / 0.25 / 0.30 / 0.40 vs contract always: −0.5 to −0.8%/yr (fractional); at $60,000 the best threshold +0.10%, the rest negative; at $100,000 all negative
B-RANKvalidated

Intraday rules on index futures, at every timeframe

the daytime session has nothing to give, and bar size does not change that
failed on its ownTested as a strategy on its own, after trading costs.

Every intraday rule we could name, on thirteen years of minute data. Nothing survives costs at any timeframe.

read the autopsy

The day-trading canon on micro Nasdaq futures, tested on thirteen years of one-minute bars, 2013–2026, a shorter window because those are the only years the broker keeps minute bars — and again on two-, three- and five-minute bars: opening-range breakouts with 5-, 15- and 30-minute ranges, with and without a stop; the first half hour’s direction held to the close; the last hour in the direction of the day; fading two and three sigma from VWAP; a 9/21 moving-average crossover. Every position is flat by the close, every trade pays a micro contract’s commission and two ticks of slippage. The breakouts lose on every timeframe and every range. The VWAP fades lose. The crossover loses badly. Two things survive as positive numbers — the opening half hour’s direction held to the close, and the last hour with the day’s trend — and both earn less than half the cash rate at a Sharpe near half, with the first one owing its whole record to 2018–2021. Changing the bar size from one minute to five flips the sign on the breakouts, which is what noise looks like. The daytime session itself earned 3% a year over the period; the return of this index arrives overnight, and a rule that is always flat overnight is fishing in the empty half.

13 years, 4 bar sizes, 9 rules · best: opening-30-min momentum 4.1%/yr, Sharpe 0.51, eras 0.23 / 1.00 / 0.32 · last-hour momentum 2.4%/yr, Sharpe 0.42 · every breakout −0.1 to −0.6 Sharpe · VWAP fades and EMA trend negative · RTH buy-and-hold 3.1%/yr, Sharpe 0.33
B-RANKmeasured

VIX-futures term-structure carry

the curve has nothing to say that the premium did not already
failed on its ownTested as a strategy on its own, after trading costs.

Selling volatility futures for the roll yield. Every filter we tried was worse than no filter, and the plain version loses to the index.

read the autopsy

Sell the front VIX future and collect the roll-down, and use the shape of the curve to know when: only in contango, only when contango is steep, only when the future sits above spot, switch long in backwardation, stand aside when the VIX is high. Thirteen years of the exchange’s own settlement prices, 2013–2026, a shorter window because the exchange publishes nothing earlier, front and second month rolled at expiry, 15% of the account in contract notional, after costs. Every curve-based filter earns less than simply being short all the time, and short all the time earns a Sharpe of 0.47 while owning the S&P earned 0.86 over the same years. The size dial changes nothing but the drawdown: at 15% of the account the strategy lost 17% in a week in February 2018 and 23% in a month in March 2020. One contract is $17,000 of notional, so a $25,000 account cannot hold a fraction of one anyway.

short front always 5.6%/yr, −24.5% DD, Sharpe 0.47 · contango filter 0.40, steep-contango 0.19, above-spot 0.30, curve switch 0.12, VIX<20 filter 0.02 · Feb-2018 −17%, Mar-2020 −23% · S&P buy-and-hold 0.86 · contango on 85% of days, mean 5.8%
B-RANKvalidated

Diversified trend following on futures

it loses to simply owning the basket
failed on its ownTested as a strategy on its own, after trading costs.

A diversified trend-following book on futures, the classic hedge-fund product. It earns less per unit of risk than simply holding the assets.

read the autopsy

The managed-futures classic: sixteen liquid futures — four stock indexes, two Treasury contracts, gold, silver, copper, three grains, the euro, coffee, sugar, cattle — each held long or short by the sign of its own past twelve months, each sized to the same risk, the whole book scaled to 15% volatility, rebalanced monthly, after commissions and slippage, 2001–2026 (a shorter window, because continuous futures histories begin in late 2000). Returns are the matching total-return funds minus bills where a clean one exists, so nobody is fooled by roll gaps. It is real: random signs on the same positions beat it in only 2% of draws. It is also worse than doing nothing clever: the same basket held long only at the same risk earns a higher Sharpe in every era, and plain 60/40 beats it too. Shorter lookbacks are worse still, so twelve months is a ridge, not a plateau. Breakouts and cross-sectional versions land in the same place. And a diversified futures book cannot be built small: bonds and grains have no micro contracts, so a $25,000 account can express under a third of the risk budget.

TSMOM 12m Sharpe 0.38 net (5.7%/yr, −32% DD) · long-only risk parity 0.58 · 60/40 0.51 · Donchian-100 0.40 · cross-sectional 0.21 · 6-month lookback −0.03 · placebo beaten 98% · era Sharpe 0.34 / 0.31 / 0.48 · $25k expresses 31% of the book in micros
A-RANKmeasured

Owning index futures overnight only

removing the bad half of the day is not the same as an edge
failed on its ownTested as a strategy on its own, after trading costs.

Holding index futures only overnight, to catch the overnight drift. There is a drift; after costs it is not worth owning.

OVERNIGHT ONLY, EDGE AT EQUAL VOLQQQ+0.90SPY-0.70
read the autopsy

Most of the Nasdaq’s return arrives while the market is closed and the daytime session has lost money on net for a quarter century, so hold micro futures from the close to the next open and sit in cash all day; futures make it possible without day-trading rules and with the kinder tax treatment. Tested on the fund’s own session bars as the proxy, 2001–2026 (a shorter window, because the fund’s own bars begin in 2000), one round trip a night, commissions and two ticks of slippage each way. The overnight leg alone does beat buy-and-hold on a risk-adjusted basis, but by a hair: at the same volatility it adds under a point a year and 0.07 of Sharpe, which is inside the noise of one era. On the S&P it is negative at equal volatility, so the two roots disagree. Sizing the overnight leg to 20% volatility lifts the return but hands back a −16.6% night, because the one thing this strategy holds is the gap. Two hundred and fifty round trips a year to buy a sliver.

QQQ overnight Sharpe 0.55 vs buy-and-hold at equal vol 0.48 (+0.9%/yr) · SPY −0.7% at equal vol · vol-targeted overnight 14.8%/yr, −39% DD, worst night −16.6% · intraday leg Sharpe 0.07 · era Sharpe 0.20 / 0.78 / 0.66
B-RANKmeasured

Turn-of-the-month on index futures

a real seasonal that is too small to be a strategy
failed on its ownTested as a strategy on its own, after trading costs.
read the autopsy

Own the S&P only for the last four and first three trading days of each month, flat the rest. The old effect is still there on 2001–2026: those seven days earn most of the month’s return at two thirds of the volatility, and the other fourteen days are worse in every era but one. But the whole strategy earns 4.6% a year at a Sharpe of 0.42, sitting in cash two thirds of the time, which is less than the engine earns while also being in the market. As a filter on when the engine adds exposure it would move a handful of trades a year by a few days. Not wrong. Not worth a rule.

turn-of-month days 4.6%/yr, Sharpe 0.42, −34% DD · other days 3.5%/yr, Sharpe 0.22, −50% DD · era Sharpe 0.28 / 0.47 / 0.51
B-RANKvalidated

Short-volatility carry

the famous one, so it gets a grave and not a footnote
failed on its ownTested as a strategy on its own, after trading costs.
read the autopsy

Sell volatility and collect the roll: the inverse-VIX fund since 2011 as the benchmark for every ‘harvest the premium’ idea. It compounded at 11% a year with a 69% volatility and then lost five sixths of itself on one evening in February 2018, a −95% drawdown it has never recovered. Sharpe 0.15 over its life. It is here because three graves above it were milder versions of the same bet, and because anyone who wants a futures strategy that pays every month should look at this line first.

SVXY since Oct-2011: 11%/yr, vol 69%, Sharpe 0.15, worst drawdown −95%, worst day −83% (5 Feb 2018)
B-RANKmeasured

The S&P 500 version of the futures engine

the futures edge needs an index that thrashes
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The futures engine works on the Nasdaq. On the S&P 500 it does not: the edge lives in a high-volatility root, and the S&P is not one.

read the autopsy

If micro Nasdaq futures beat the 2× fund for the engine, micro S&P futures should do the same for the S&P sleeve, and the contract is smaller, so the rung would open lower. Same full-history test as the Nasdaq study: the full 1999–2026 cycle after tax, the S&P fund pair with the 2× fund modelled before 2006, one micro contract held at today’s size plus the fund remainder, both dials, the contract both never outgrown and outgrown from a $25,000 start, at five account sizes. It loses at every size and on both dials, by a third to half a point a year, and fails the exposure-matched control at $25,000. The mechanism explains why: the Nasdaq edge is mostly the 2× fund’s daily-reset decay in violent markets plus the tax treatment on a rule that trades often. The S&P is calm enough that its 2× fund barely decays, the tax paid is the same either way, and constant-notional futures simply forgo the fund’s compounding in steady years. Only the dot-com era was positive. Futures earn their keep on the index that thrashes; on the one that does not, the fund is the cheaper vehicle.

after tax, contract never outgrown: $25k −0.40%, $40k −0.47%, $80k −0.42% vs the funds · matched control at $25k −0.33% · 2009–16 −0.7 to −1.7% · tax paid $71k funds vs $68k futures
B-RANKmeasured

Capping the premium arm when the exit calls a bear

the loss arrives before the confirmation
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A version that capped the option arm whenever the exit called a bear. The cap came too late every time; the damage was already done.

read the autopsy

If the put-selling arm loses in long bears, let the core’s trend exit switch it off, or cap it at single size, once a bear is confirmed. Both roots, SELECT and ULTRA alike, the same full-cycle ladder as the graves above: switching off removes the 2008 losses and also the recoveries the arm was built to sell into, so the Nasdaq ladder earns less than a constant of the same average size; capping at single size does the same with more drawdown. The reason is timing. The exit confirms a bear only after two month-end closes, and the arm’s worst trade of all, sold in February 2001, was entered four weeks before the exit fired; the 2008 losses were sold into a bear the exit had already called. Shuffling the arm’s results onto random dates still beats the switched-off version on return. A slow exit cannot protect a fast loss, and the fast loss is the one this arm sells.

QQQ off-when-exited 0.85%/yr vs a size-matched constant 1.92% · capped 1.34% vs 2.27% · as an overlay, random timing beats it in 68% of draws on SELECT and on ULTRA
A-RANKmeasured

The premium overlay on the exit engines

its losses land in the core’s worst windows
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We tried an option-selling overlay on top of the exit engines. The losses arrive before the exit can react, so the overlay adds risk the engine cannot see.

read the autopsy

The plan for larger accounts was to run the put-selling arm on top of SELECT or ULTRA, sized at a fifth of the account, so the premium adds to the core. On the full cycle, with the same modelled-then-real put prices as the grave above, it adds about half a point a year to either engine and costs SELECT four points of drawdown. Dialling SELECT itself up to that drawdown earns 1.3 points more than the overlay does. Shuffling the same trade results onto random dates beats the real timing two times in three on return and nine in ten on drawdown: the arm loses in exactly the months the core loses, 2001 and 2008 above all, so its timing is worse than chance. Selling one put a month at fixed size keeps the return and still loses the drawdown contest to chance. And at today’s prices one micro Nasdaq index put is 1.1 times a $25,000 account, so the size tested here cannot be built below roughly $140,000. The rung comes off the ladder.

SELECT +0.57% / −4.0 DD, matched core −1.30% · ULTRA +0.59%, random timing beats it 68% / 61% · 2001 −7.3%, 2008 −6.0% of the account at one-fifth size
B-RANKmeasured

Selling more insurance after the fire, into a long bear

the fire does not always go out
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The idea was to sell more insurance right after a crash, when it pays best. In a long bear the crash keeps coming, and the seller keeps paying.

read the autopsy

The premium overlay scoped for larger accounts sells a monthly index put 5% below the market and sizes it up as realized volatility rises — sell more insurance after the panic, when it is richest. It passed its controls on 2012–2026 because those are the only years with full option quotes. That window holds no bear market longer than three months. To see a long one, the premium was modelled for 1993–2011 from the S&P and Nasdaq volatility indexes, fitted where real quotes exist (it reproduces real trade results to a correlation of 0.998), while every trade still settles on real prices. The rule sizes to its 2.0x maximum through 2001 and again through late 2008, and the market keeps falling: on the Nasdaq it loses 37% of its collateral in 2001 and 30% in 2008; on the S&P, 36% in 2008. Its worst drawdown is roughly double that of simply selling one put a month, at every premium assumption tried, from a quarter below the model to twice it. After the fire the premium is richest; in a long bear the fire is still burning. B because the early premiums are modelled — the losses are settlements, which are not.

2001–26 QQQ: sized-up 1.2%/yr, −42% DD vs one-a-month 2.5%/yr, −26% · 1993–26 SPY: 2.7% / −47% vs 3.2% / −28% · 2008: −30% Nasdaq, −36% S&P
B-RANKvalidated

Timing the contributions

waiting is a bear-market bet in disguise
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We asked whether the timing of monthly deposits matters. It does not: money added on any day of the month ends up in the same place.

FINAL WEALTH VS DEPOSIT NOW, %buy the dip-1.40wait: de-risk-11.40wait: exit-57.80random delay-5.80
read the autopsy

A small account grows mostly by what its owner adds, so: should new money wait for a dip? Five thousand dollars plus $500 a month, run through ULTRA and SELECT from 1999 and from 2009, with four ways of holding the money back — until the index is 5% off its 20-day high, until the rule itself has de-risked, until the trend exit has fired, or for a random delay of up to four months as the control. Investing each contribution the day it arrives beat every rule in the modern window: the dip rule by 1.4 points of final wealth, the de-risked rule by 11, the exit rule by 58, and all twenty random delays. Over the full window the dip rule shows a 6% gain, but so does a random delay, which beats immediate investing in two draws of three — that is the 2000 crash sitting at the start of the sample, not a rule.

2009–26 final wealth vs immediate: dip −1.4% · wait-for-de-risk −11.4% · wait-for-exit −57.8% · random delay −5.8% (immediate beats 20 of 20)
C-RANKmeasured

Options at five thousand dollars

the contract is bigger than the account
small account or this brokerTested as a $5,000 cash account. Every option trade it could place was bigger than the account.
CONTRACT SIZE / ACCOUNT, XQQQ put+13.00XND put+5.50QLD put+1.70
read the autopsy

A cash account can do three things with options: buy them, sell calls against a hundred shares it already owns, or sell puts with the whole strike held in cash. Priced off the live order book on 2 September 2026, every one of those on the Nasdaq-100 or the S&P 500 is between 1.7 and 14 times a $5,000 account. One cash-secured QQQ put needs $67,200; the micro Nasdaq index put $27,600, and it traded one contract that day at a 12% spread; the 2x fund’s put $8,400 at a 30% spread. The two underlyings cheap enough to fit — a $30 large-cap fund and the inverse Nasdaq fund — had no bid at all. An option is a hundred shares or nothing, and the rule sizes in single shares. Arithmetic rather than a backtest, hence the B: no test can move a contract multiplier.

QQQ put $67,200 (13x) · XND put $27,600 (5.5x, 12% spread) · QLD put $8,400 (30% spread) · SCHX / PSQ puts: no bid
B-RANKvalidated

The exit at weekly cadence

faster checks re-admit the churn
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
WEEKLY EXIT, PP/YRSELECT k=2-1.95
read the autopsy

The monthly trend exit waits for two month-end closes below its band. Checking weekly with a matching number of weekly closes looked like the same rule with less lag. It loses at every setting: two weekly closes give up two points a year and five points of drawdown; four give up a point; eight match the return but add drawdown on each engine tried. A faster clock re-admits whipsaws faster than the confirmation can remove them. The month-end close is the right resolution for a trend exit, and the rule ships at that clock.

SELECT weekly k=2 −1.95% / −5.2 DD · k=8 −2.3 DD · ULTRA k=8 −4.0 DD
B-RANKvalidated

A faster volatility estimator, with the exit

a dial that spends the worst day
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

Before the exit existed, faster volatility estimators lost outright: quicker de-risking bought quicker re-risking and the whipsaws cost more than the lag saved. With the exit now handling the extended bears, the question was worth asking again, and the answer changed shape rather than sign. A ten-day estimator adds about half a point a year and a hair of Sharpe — and pays for it with a worse drawdown and a worst single day of −7.5% against −5.9%. It is not an edge; it is the tail budget being spent. The twenty-day window stays.

rv10 +0.59%, DD −1.6, worst day −7.5 vs −5.9 · rv15 the same shape
B-RANKmeasured

Bonds or gold while the exit is out

the waiting room should be boring
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The exit sits in cash about a sixth of the time, mostly in bear markets. Long Treasuries and gold are the textbook things to hold then. Measured from 2005, where the funds exist, both lose to plain bills: bonds give up 0.7 points a year and fourteen points of drawdown, gold adds twenty points of drawdown for a rounding error of return, and even short bills trail. The exit’s cash is not a bet; it is the absence of one, and every asset with its own drawdown breaks that.

TLT −0.71% / −14.4 DD · GLD +0.14% / −20.0 DD · SHY −0.32% (2005–26)
C-RANKreviewed

Selling puts only while the exit is out

eight trades in fourteen years
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The premium overlay sizes up when volatility is high, and the exit steps the core into cash in exactly those markets — so the obvious refinement was to sell puts only while the engine is out and the collateral is free. There is nothing to test: over fourteen years of real option chains the core was out at the moment of entry for eight of a hundred and forty-five trades, and those eight earned a third of what the others did. The overlay itself passed its timing placebo on the exit core; the gating is a non-result and is filed as one.

8 of 145 entries while out · their return +0.19% vs +0.62% · not a sample
B-RANKvalidated

Session-split leverage

blind is where the money is
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The cleanest idea anyone brought this year: keep full exposure while the market is open and observable, hold less through the night, when nothing can be done about a gap. No forecasting, just a refusal to rent leverage during the blind window. The decomposition killed it before the contest did. Since 1999 all of the NASDAQ fund’s return arrived overnight — the open-to-close leg lost money over twenty-seven years — and three-quarters of the damage on its worst days happened while the market was open. The blind window is the paid one. Every version that lightened overnight exposure was beaten by simply running the standard dial, and the mirror — full leverage overnight, less by day — only reinvented that dial with three hundred extra trades a year.

overnight +13.9%/yr vs intraday −2.8%/yr · 73% of bad-day loss is intraday · every λ dominated by the flat dial
B-RANKvalidated

Re-entry confirmation

patience only works on the way out
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The exit waits for two consecutive monthly closes below its band before leaving, and that removed the whipsaws. The mirror — wait for two closes above before returning — was the obvious next test. It costs return on each engine it was tried on, with no drawdown benefit at all: the account just misses the first month of every recovery. Whipsaws are exit-side events; patience on the way in is only lateness.

QQQ engine −0.25%, STEADY −0.33% at k=2 · drawdown unchanged · Sharpe down on both
B-RANKvalidated

A VIX veto on the exit

the exit it targeted did not exist
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

Under the plain exit the ugliest episode was April 2020: out a month after the low, back in a month later, a 14% rebound missed. A veto that ignores the exit when volatility is already extreme seemed to target it exactly. But the two-close rule had already removed that exit — April 2020 was a single close below the band. Under the shipped rule the only high-volatility exit left is November 2000, the one worth +192%, and every veto threshold that fired vetoed that instead.

thresholds 40–60% veto the dot-com exit: −0.75%, drawdown −31 to −38 · at 80% it never fires
B-RANKmeasured

Hysteresis on the conditional sleeve

the exit already does its job
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The conditional sleeve held a trend fund only while the rule’s target weight sat below a threshold — one magic number, most of its gain from 2022. A band around the threshold might have made it robust. Twelve threshold-and-band cells; all twelve lost to plain bills, and the best beat fewer than a third of random switch schedules. With the exit in the engine, the account is already in cash during the markets the sleeve was built to cushion. Measured on a managed-futures proxy from 2007.

12 of 12 cells below bills (−0.3 to −1.1%) · best cell beats 30% of placebos
B-RANKmeasured

The trend sleeve on an exit engine

a defence counted twice
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The trend tier beat cash on the no-exit engine, and the site said so. Measured per engine for the first time, it lost on every engine that has the exit: one to two points a year and five to nine points of drawdown, on real fund history since 2019 and a proxy since 2007. The exit sits in cash in exactly the markets the sleeve was built for; a trend fund held through an exit adds its own drawdowns on top of a defence already in place. This one changed the product: the sleeve was withdrawn from sale.

SELECT −1.95% / −9.3 DD · ULTRA −2.0 / −7.4 · STEADY −2.2 / −9.0 (DBMF, 2019–26)
B-RANKvalidated

A third exit parameter

two mechanisms pulling opposite ways
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The band and the two-close rule each passed alone. Stacking a faster weekly re-entry on top, or a different band per sleeve, looked like more of the same. The grid said otherwise: the stacked rule beat the simpler one in nine of eighteen cells, its best cell was the overfit, and its whipsaw count went back up — a faster way in re-admits the churn the patient way out had removed. The per-sleeve band moved the result by a tenth of a point, which is noise. Three parameters is where fitting starts; the rule ships with two.

stack: 9/18 cells better, whipsaws 33% → 56% · per-sleeve band +0.14%
B-RANKvalidated

VIX as a thermometer

the cap dial wearing a badge
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
WORST DAY, % (EVERY VIX RULE)cap 2.0 always-8.20VIX<12 cap 1-8.70VIX<20 cap 1-8.20mirror VIX>20-7.70
read the autopsy

The idea: cap the hottest engine lower whenever VIX is calm, since the damage in a levered book comes from the spike out of quiet. Ten thresholds, two cap levels, the mirror rule and a random-day placebo. Every VIX rule left the worst single day exactly where it was, and the risk-adjusted gain rose in lockstep with the fraction of days capped — capping on VIX just meant capping most of the time, because when VIX is low the engine is already at its cap. The mirror, capping when VIX is high, did nothing at all. The worst days themselves settled it: five of the eight came out of calm, three out of stress, and the worst of all was a band-lag day the cap could not reach. VIX and the engine’s own thermometer read the same fever at the same time.

worst day −8.2% under every rule · Sharpe tracks % of days capped, not VIX · mirror inert
B-RANKfull battery

Defined-risk put spreads

beaten by a coin flip
failed on its ownTested as a strategy on its own, after trading costs.
read the autopsy

Selling a spread instead of a naked put posts only the width, so an income strategy that normally needs a five-figure account should compress into a few hundred dollars — the one structure that could have filled the gap below the options tier. It returned more than the core and looked better on every risk-adjusted measure we had. Then we shuffled the same trades onto random expiry dates. The random version scored higher. The timing was worth nothing: the flattering number came from a 93%-win-rate stream whose losses all arrive at once, which is the oldest way there is to fool a Sharpe ratio. It also loses three times what the core loses on the core’s worst days, and the configurations that actually worked needed a bigger account than the strategy they were meant to undercut.

beat only 12% of random-timing placebos · 3× the core’s loss in its worst decile · needed $31k to undercut $25k
A-RANKvalidated

A wider re-entry band

real, and inexplicable
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
SHARPE CHANGE VS SHIPPED BANDQQQ+0.04SPY-0.02
read the autopsy

Widening the buy-back threshold beat the shipped rule on every axis at once — and survived the bootstrap and the exposure-matched control that killed a dozen other ideas. Then we asked why, twice. The re-entries it skips outperform the ones it keeps. The whipsaw rate it supposedly avoids is identical. And across eighteen assets the effect helps most exactly where this strategy works worst. An edge that passes every statistical test but cannot say what it does is a coincidence wearing a lab coat.

+0.041 Sharpe on QQQ · mechanism failed twice · negative on the S&P
B-RANKvalidated

STRAT bar combos

look-ahead
failed on its ownTested as a strategy on its own, after trading costs.

The detector silently deleted every trade that stopped out on its trigger bar — 99.4% losers. One flag flips +0.13R to −0.16R.

3.75M trades · win 31% < coin-flip floor
B-RANKvalidated

2-1-2 & 3-2-2 "consensus" combos

shared measurement error
failed on its ownTested as a strategy on its own, after trading costs.

The internet's favorite combos rank exactly where the bug flattered them. Honest ranking anti-correlates ρ −0.70.

consensus picks = honest bottom three
B-RANKvalidated

Liquidity sweep radar (LSR)

fill fiction
failed on its ownTested as a strategy on its own, after trading costs.

Entry priced at a level the trigger made unreachable — the signal needs the close that proves you missed the fill.

defect persists at every timeframe
B-RANKvalidated

Session H/L sweeps

same corpse, new outfit
failed on its ownTested as a strategy on its own, after trading costs.

Identical mechanism to LSR under a different name. Fade ≡ continuation within 0.004R.

112,311 events
A-RANKvalidated

Opening range breakout

real, but sub-cost
failed on its ownTested as a strategy on its own, after trading costs.

Carries ~+1.2bp of genuine information — smaller than the spread it must cross, at every account size and liquidity tier.

stability or profit, never both
B-RANKvalidated

ICT / SMC confluence

the strict all-five setup never occurred
only a strict version testedTested as the strict version, all five pieces lining up at once. Most traders use two or three of them; those pieces are tested one by one on their own stones: fair value gaps (76), inverse FVG (75), sweeps (79, 80), prior-day levels (70) and the level-plus-bias workflow (69).

Sweep + FVG + iFVG + HTF bias + daily levels: full 5-of-5 confluence happened zero times. More confluence tested worse.

0 occurrences in 5.9M observations
B-RANKvalidated

Fair value gaps

flat
failed on its ownTested as a strategy on its own, after trading costs.

539,314 trades of nothing. The gap fills or it doesn't; your P&L can't tell the difference.

±1bp vs 5.5bp cost
B-RANKvalidated

Inverse FVG

flat, twice
failed on its ownTested as a strategy on its own, after trading costs.

Rejected in July, re-killed in August inside the confluence stack. Every cell negative.

−0.07 to −0.11R across all cells
B-RANKvalidated

Inside bars

can't pay the toll
failed on its ownTested as a strategy on its own, after trading costs.

Statistically real (+3.2bp) and half the size of its own round-trip cost. Buying the break made it worse: the breakout subset already moved.

0.57× cost · breakout −11.2bp
B-RANKvalidated

Outside bar patterns

wrong sign
failed on its ownTested as a strategy on its own, after trading costs.

"Three outside down" — the textbook bearish pattern — resolved bullish.

+5.24bp where theory says down
B-RANKvalidated

Hammer / wick rejections

fill fiction
failed on its ownTested as a strategy on its own, after trading costs.

+27.7bp collapsed to +3.2bp with honest fills — a third of breaks gap straight through the level.

t +4.6 → +1.9 honest
B-RANKvalidated

Timeframe continuity (FTFC)

a filter, not an edge
failed on its ownTested as a strategy on its own, after trading costs.

Genuinely improves whatever it's attached to (+0.07R) — and the best base it can improve still loses.

−0.034/day t −11.5 standalone
B-RANKvalidated

Prior-day level touches

wrong forecast
failed on its ownTested as a strategy on its own, after trading costs.

The touch predicts volatility (+49bp of movement, t +41) and says nothing about direction.

direction: +0.18bp on 19,433 touches
B-RANKvalidated

Level-test + HTF bias workflow

well-powered null
failed on its ownTested as a strategy on its own, after trading costs.

Mark levels, read 4H/1H bias, fade the test: zero of 28 cells clears cost. Every cell negative net.

MDE 2.36bp — it could have seen it
Exits, filters & management
B-RANKvalidated

Trailing stop geometry

indistinguishable
failed on its ownTested as a strategy on its own, after trading costs.

65 trail/tighten configurations landed within 0.0007R of each other. All negative.

corr(train,test) +0.03
B-RANKvalidated

Giveback / exit tuning

null
failed on its ownTested as a strategy on its own, after trading costs.

Exit geometry cannot manufacture entry edge. The stop's denomination was the only real defect — and that was a bug fix, not a strategy.

62% of stop-outs: underlying <1% against
B-RANKvalidated

ATR-scaled stops

effects cancel
failed on its ownTested as a strategy on its own, after trading costs.

Wider stops on volatile names lose more when they still stop. Fixed vs scaled: a wash.

−1.51 vs −1.58bp
B-RANKvalidated

ATR conviction exit

exposure reducer
failed on its ownTested as a strategy on its own, after trading costs.

Looked like the find of the month — until tested on a winning base, which it destroyed too. It shrinks everything toward zero.

7 of 7 families pulled toward 0
B-RANKvalidated

RVOL confirmation

backwards
failed on its ownTested as a strategy on its own, after trading costs.

"Strong breakouts need volume." High-volume breakouts did worse in both liquidity tiers.

+1.33 → −1.11bp with the filter
B-RANKvalidated

VWAP / EMA20 alignment

decorative
failed on its ownTested as a strategy on its own, after trading costs.

Trading "with" the anchor changed nothing that mattered.

within noise of baseline
B-RANKvalidated

Volume profile (VAL/POC)

era-unstable
failed on its ownTested as a strategy on its own, after trading costs.

Below-value-area looked like the first real pass — then the continuous version flipped sign across eras. Full-sample coefficient: exactly zero.

−3.96 → +4.52 across eras
B-RANKvalidated

Order-flow proxy delta

too small
failed on its ownTested as a strategy on its own, after trading costs.

Sign-stable in both eras and worth ~1bp against 5.5bp of cost.

real order flow: a $100-500/mo bet we declined
Mean reversion, momentum & the book
B-RANKvalidated

Gap fades (index scalping)

placebo won
failed on its ownTested as a strategy on its own, after trading costs.

Random entries on the same days, same side, same exits matched the signal. The timing added nothing.

day-clustered t +1.02
B-RANKvalidated

Overnight gap → intraday

decayed
failed on its ownTested as a strategy on its own, after trading costs.

Perfectly monotone once (t −10.2), then 16.3 → 8.3 → 5.0 → 2.7bp by era. The cash-tradeable version extrapolates negative.

long-only: −8.96bp in 2018-21
B-RANKvalidated

Residual momentum

it reverses
failed on its ownTested as a strategy on its own, after trading costs.

Top-decile residual strength predicted negative forward returns. Killed by its own pre-registered criterion in one day.

−8.5bp t −2.3 · cost 2-4× effect
B-RANKvalidated

Residual breakout (20d high)

underperforms
failed on its ownTested as a strategy on its own, after trading costs.

Third independent confirmation that residual strength reverses.

−4.30bp t −2.58, outside a 200-draw null
B-RANKvalidated

RSI2 / SR20 dip-buying book

beta in a costume
failed on its ownTested as a strategy on its own, after trading costs.

158% of its P&L was market exposure; the stock-picking contributed negative alpha. SPY with 7,803 extra steps.

idio −0.028R t −8.6
B-RANKvalidated

Breadth veto

mechanism ≠ evidence
failed on its ownTested as a strategy on its own, after trading costs.

Expectancy is flat across signal-breadth. The scary cluster was 26 days of noise; the veto made things worse.

corr +0.014 on 3,758 days
B-RANKvalidated

Capital allocators

nothing beats random
failed on its ownTested as a strategy on its own, after trading costs.

No ranking beat random over 20 seeds. The real constraint was same-day correlation: a dip-buyer buys one dip N times.

7.8 signals/day, N_eff ≈ 3
B-RANKvalidated

Mean-reversion short (idio)

adverse selection
small account or this brokerTested as a small account, which ends up taking the worse half of the trades. The edge itself was real.

The edge was real in every era — and a small account takes the worse half of it. Simulated honestly: $44 profit in 17 years.

taken +0.015R vs skipped +0.032R
Options & published anomalies
B-RANKvalidated

Option bid-ask spread

one crisis wearing a disguise
failed on its ownTested as a strategy on its own, after trading costs.

What market makers charge to take the other side — the last untouched family in the options archive. Scored the best warning-power we had ever measured. Then the episode ledger: twenty-four of the forty-three warnings were the 2020 crash. Strip that quarter and it is noise.

lift 2.43x · ex-COVID 1.21x · 2022–26 zero
B-RANKvalidated

Option volume

free data did it better
failed on its ownTested as a strategy on its own, after trading costs.

Options traded, not options held. Genuinely predictive, and then plain SPY share volume — free — scored the same. Paying for the options feed bought nothing the tape did not already say.

5.07x vs share volume 5.11x · precision 6% vs 25% needed
B-RANKmeasured

Short-selling volume

nothing there
failed on its ownTested as a strategy on its own, after trading costs.

Public FINRA data on how much of each day's volume was short. Screened against frozen events on both symbols and both targets. Two of the four scores were worse than random.

0.36x–1.05x · none significant
B-RANKvalidated

30-day variance premium

decay, not cost
failed on its ownTested as a strategy on its own, after trading costs.

Seller Sharpe +0.74 → +0.09 → −0.12 across eras. Friction was only 0.85% of credit — the edge itself left in 2012.

first cluster to die of decay
C-RANKvalidated

0DTE ATM premium

structurally unreachable
small account or this brokerTested as a small account at our broker. The part that pays needs the full index contract, about 10 times too big for it. Not tested on a large account.

The premium lives in the OTM put wing; the tradeable wrapper doesn't exist at this broker and size.

SPY physical settlement · SPXW 10× too big

Retested on Robinhood (R-148, 1 October 2026). The door that was shut has opened: Robinhood trades the small cash-settled S&P option this needed, so we ran the afternoon trade on its own real minute-by-minute prices, every day it could be traded since 2016: 1,655 trades. It lost $30 a trade. On the clean, liquid days from 2023 on (834 of them) the premium was still there at the midpoint, about $4 a trade, but getting in costs about $8 in the gap between buy and sell prices plus $2.60 in fees, so it lost $6.65 a trade. A real premium, smaller than the cost of collecting it.

B-RANKvalidated

PEAD

dead since 2006
failed on its ownTested as a strategy on its own, after trading costs.

Post-earnings drift was arbitraged away while smartphones were new.

two decades gone
B-RANKvalidated

Index-effect

arbitraged
failed on its ownTested as a strategy on its own, after trading costs.

S&P inclusion pop: 7.4% in the old papers, 0.3% now.

7.4% → 0.3%
B-RANKvalidated

Overnight premium

gone
failed on its ownTested as a strategy on its own, after trading costs.

The close-to-open drift flatlined in 2021.

≈0 for four years
B-RANKvalidated

Dispersion trades

insignificant
failed on its ownTested as a strategy on its own, after trading costs.

−0.19 vol points, p 0.155. Not even close after costs.

n.s.
B-RANKvalidated

VIX carry

mispriced now
failed on its ownTested as a strategy on its own, after trading costs.

Currently below its historical average while the tail risk is unchanged. Selling it here is picking up dimes at a discount.

carry < average, tail intact
B-RANKvalidated

Betting against beta

microcap mirage
failed on its ownTested as a strategy on its own, after trading costs.

$1.05 of every $1 of profit came from the bottom 1% of market caps — positions that cannot actually be filled.

A-tier retraction in a costume
B-RANKvalidated

Scalping (1-3m timeframes)

cost regressivity
small account or this brokerTested as a small account, where a $50 trade pays the $0.35 minimum fee (1.4% there and back). On a bigger account that fee is a much smaller share.

$0.35 minimum commission is 140bp on a $50 position. The account size makes the timeframe illegal, economically.

cost floor > any measured edge
C-RANKvalidated

The 5%/month target itself

mathematically impossible
failed on its ownTested as a strategy on its own, after trading costs.

Requires Sharpe ≥ 1.044 at optimal leverage. Medallion's gross was 4.3%/mo. The goal was the bug.

g = rf + Sσ − σ²/2 has a ceiling
Volatility forecasting — better σ, same strategy
A-RANKvalidated

Volatility breadth beneath the index

true, and unusable
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

How many parts of the market are getting choppy while the index still looks calm. Strong, real, survives every control — and neither way of acting on it works. Slowing purchases never triggers; lowering the ceiling loses to simply lowering the ceiling all the time.

2.86x warning power · every conversion negative
B-RANKvalidated

Damage without volatility

the premise was backwards
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The idea that a slow bleed which never rattles the volatility reading is a warning. It is the opposite: those stretches were followed by better returns than average, at exactly the usual odds.

+1.71% vs +1.38% forward · 2022 unchanged
B-RANKvalidated

Reacting to volatility faster

speed cuts both ways
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A shorter window reaches safety in eleven sessions instead of forty-eight. It also loses more in 2008 and more in 2022, because whatever de-risks quickly re-risks quickly, and the round trip costs more than the delay saves.

2008 −28.8% vs −27.2% · double the trading
B-RANKvalidated

Fast exits with slow entries

it was only holding less
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The obvious fix to the above: quick to sell, unchanged to buy. Looked excellent — five points less drawdown. Then a version that simply holds a little less on average matched it, without the extra trading.

fails its matched control at 3.3x turnover
A-RANKvalidated

Range estimators (Parkinson, Garman-Klass)

forecast better, traded worse
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Both beat close-to-close at predicting next month’s volatility — because they read the intraday range and ignore the overnight jump. That is precisely what makes them blind to gap risk: they read calm into a market that gapped, and lever into it.

+6–7 points of forecast R² · deeper drawdowns
B-RANKvalidated

Yang-Zhang

one episode
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The one range estimator that does see overnight gaps, and it did improve results — until the gain was traced to a single stress episode. Failed both the era split and the bootstrap.

fails leave-one-crisis-out
A-RANKvalidated

HAR-RV (daily / weekly / monthly)

one episode
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The best full-sample arm ever tested here. Remove the dot-com bust and it loses to the simpler rule it was meant to replace. A single episode supplied the entire advantage.

Sharpe 0.86 full-sample · ex-dot-com CAGR 14.7% vs 17.4%
B-RANKvalidated

Credit spreads & sector dispersion

no incremental information
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Both are widely cited as early warnings of equity stress. Neither told the rule anything that twenty days of realised volatility had not already said.

no improvement over the control
Attempts to improve the survivor
B-RANKmeasured

A cheaper share price

backwards where it should have helped
small account or this brokerTested as a $4,000 account. It only helped on big accounts, around $100,000.

Small accounts cannot buy fractions of a share, so a lower-priced fund should track the target more closely. At four thousand dollars it was worse, not better. The benefit only appears on large accounts that never needed it.

−0.61% at $4k · +0.70% at $100k
B-RANKvalidated

Waiting longer to buy back in

time carries nothing size does not
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Making the re-entry hurdle shrink the longer you have been out. Three different speeds, all indistinguishable from doing nothing, with a fifth more trading.

18.43–18.53 vs 18.46 control
B-RANKvalidated

A bear-market mode

the rule got there first
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A separate, deliberately slow state for grinding bear markets — months of evidence before it may act. It changed nothing at all: by the time months confirm a bear, volatility has already cut the position below anything the mode would impose.

identical return, drawdown and trade count
B-RANKvalidated

A floor after long drawdowns

trend-following in disguise
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Capping exposure once the account has been underwater a long time. The control that holds the same average position beat it comfortably, and it was worse in two eras out of three.

−0.04 Sharpe vs matched · era-negative
B-RANKvalidated

Extra de-risking in high volatility

one episode
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Cutting exposure further once volatility passes 30% beat the live rule on CAGR, Sharpe and drawdown, and did so in all 25 parameter settings tested. Then the episode ledger: 34.3 of the 35.3 point advantage came from the single 1999–2003 episode. Across the other sixteen episodes it is worth +1.0% in total, and through the calm 2010s it is strictly worse. A parameter sweep cannot detect that — every cell was measuring a drawdown set by the same crisis.

median episode −0.15% · 7 of 17 better
B-RANKvalidated

Leveraged inverse sleeves

beaten by subtraction
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Holding an inverse position during stress lifted Sharpe from 0.74 to 0.80. But simply cutting the target weight by the same amount of exposure beat it on every measure, with fewer trades and no second instrument. The short only looked good because nothing was racing it.

0.80 Sharpe vs 0.87 for plain subtraction
B-RANKvalidated

A 3× sleeve above the cap

leverage, not edge
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Extending the ceiling to 3.0 in the calmest regimes raised CAGR and lowered Sharpe, with drawdown unchanged and turnover up 2.3×. Buying return with risk-adjusted return is not a discovery; it is a dial that already exists.

CAGR 13.4→14.2% · Sharpe 0.74→0.72
B-RANKmeasured

Covered-call income overlays

cancels itself
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

High-income Nasdaq funds cut volatility by about 3 points and cost about 12 points of return. The deeper problem is structural: this rule sizes positions on volatility, so a lower-volatility fund receives a larger weight. The target levers the fund back up, restoring the risk the call selling removed while keeping the capped upside.

you pay for the insurance, then undo it
B-RANKmeasured

A finer share grid

commission minimum
small account or this brokerTested as a small account, about $3,300, paying a minimum fee on every order. On a bigger account the fees matter far less.

A cheaper share-price proxy for the same index cuts the position grid from 21.7% of a small account to 8.9% and tracks the target more closely. It also crosses the rebalance band more often. At a per-order minimum, the extra trades cost more than the tracking error they fix.

19.64% vs 19.72% at a $3,270 account
A-RANKvalidated

Volatility-instability state (fragile calm)

real forecast, no trade
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Low volatility that is itself unstable transitions to high volatility 1.7× more often — replicated on two indices, controlling for the level. Every way of trading it lost: the alarm is wrong nine times in ten, and the rule already de-risks within days of volatility actually rising. Knowing slightly earlier bought nothing, again.

forecast p 0.019/0.004 · 9 of 9 trading configs worse
B-RANKmeasured

Account-size band formula

flat plateau
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Theory says a small account’s no-trade band should widen with the cube root of its cost burden — about 2× at $3,300. Measured: the empirical optimum stays at 1×, and every width from half to 1.5× lands within noise of every other. The one robust finding is that very wide bands hurt everywhere. There is no dial here.

48 windows per cell · 0.5–1.5× spread < 0.2%
B-RANKmeasured

Expiration-week volatility release

too small to act on
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

If dealer hedging suppresses volatility into monthly option expiry, the measure the strategy reads understates risk on a known calendar. The release exists with the predicted sign — and it is worth 0.85 volatility points against a 19-point base, which moves the target by less than the smallest change the rule is allowed to act on. Real, maybe; actionable, never.

321 expirations vs 1,061 placebo Fridays · p 0.274
Other assets & rotation
B-RANKvalidated

Spreading the rule across markets

the same weather everywhere
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Running the identical rule on four US indices and averaging. The delays do not cancel out because the volatility arrives everywhere at once, and averaging drags you toward the weaker markets.

−5.1%/yr · streams correlate 0.82–0.90
B-RANKmeasured

Bitcoin

the volatility is the return
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Volatility targeting halves bitcoin’s drawdown — and halves its return. The Sharpe gain is +0.14 measured from 2015 and +0.03 from 2021. High-volatility states pay 28 points more than calm ones, so cutting exposure there sells the advance. The levered sleeve can never engage: it needs volatility under 20%, and bitcoin’s tenth percentile is 27%.

vol-return spread +28% · persistence 0.38
B-RANKvalidated

Higher-volatility assets generally

the premise is backwards
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Across 25 assets, what predicts whether volatility targeting helps is not how volatile something is, but whether its high-volatility states pay. The correlation between volatility and benefit is negative. In Solana, Dogecoin and GameStop the forward return in high-volatility states exceeds the calm-state return by 313, 218 and 132 points — there, the method systematically sells the advance.

spread t −3.72 · corr(benefit, volatility) −0.44
B-RANKmeasured

Concentrated semiconductors

lateral move
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Semis have the highest volatility persistence of anything tested, and genuinely suit the method. They also deliver the same CAGR as the Nasdaq since 2000 at a lower Sharpe, with far more concentration risk. A better fit for our full-history test is not the same as a better destination.

persistence 0.75 · Sharpe 0.64 vs 0.74
B-RANKmeasured

Sector rotation

fights the method
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Rotating by momentum, or into defensives during stress, degrades the two properties the method depends on: volatility persistence falls and high-volatility states flip from paying less to paying more. Changing what you hold puts regime breaks into the volatility series you are trying to forecast. The defensive version failed at its own stated job, deepening the worst drawdown rather than reducing it.

defensive maxDD −56.1% vs control −50.9%
C-RANKreviewed

The chart-indicator shelf

wrong space
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Thirty-five of the most popular published indicators were reviewed for anything this strategy could use. None qualified. They operate on chart bars, while every real constraint here — whole-share grids, two-leg composition, settled cash, per-order minimums — lives in the account, which a charting language cannot see.

35 reviewed · 0 usable
The sleeve auditions
A-RANKmeasured

Timing the trend sleeve

the diagnosis was right, the cure was not
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Managed futures stop protecting you sometimes, and it is measurable when. But they are in that state seventy percent of the time, so stepping aside forfeits the yield that pays for the whole sleeve.

correct signal · −1%/yr to act on it
B-RANKvalidated

Spreading the idle cash around

concentration won
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Splitting the waiting cash across trend, bills, treasuries and gold instead of one diversifier. Every blend landed between the two things it was mixing. The gold blends only looked good because gold returned ten percent a year.

19 years · 3 crises · no mixture beat trend alone
B-RANKvalidated

Half the sleeve in the index

it undid the engine's own work
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

We shipped this tier and retired it. Parking half the waiting cash in the same index the rule had just sold meant the sleeve rose and fell with the market at the worst moments — buying back with one hand what the other had sold.

on real funds: worse return, worse risk, worse 2022
B-RANKmeasured

Bitcoin-treasury preferreds (STRC, SATA)

wrong-way volatility
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A ~10% coupon from leveraged bitcoin-treasury issuers, marketed to trade like cash. The cash sleeve is largest during equity crashes — and these fall on exactly those days, with a −24% drawdown in their first year, which contained no crisis at all. The yield is crash insurance sold on the same tail the strategy is long.

−0.7% avg on QQQ’s worst days · no winter ever survived
B-RANKmeasured

Covered-call funds as the sleeve

beaten by what they hold
failed on its ownTested as a strategy on its own, after trading costs.

SPYI, XYLD, QYLD, JEPI and kin — equity beta with the upside sold off. The decisive test was the fair control: at like-for-like, the wrapper lost to plainly holding the index by over a point a year at identical risk, and the decade view is worse everywhere. The oldest of the family fell 43% through the financial crisis.

10+ years · dominated in every window tested
A-RANKmeasured

Short-volatility carry (SVOL)

the siren
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

A 21% distribution yield, and it produced the highest portfolio return of any sleeve tested — alongside the worst drawdown, because the yield is the premium for insuring the exact crashes the strategy exists to survive. The single most seductive wrong answer in the contest.

+3.3% CAGR for +9.5% of drawdown · refused
B-RANKmeasured

High-yield credit & preferreds (HYG, SRLN, PFF)

falls with the tide
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Junk bonds, bank loans and preferred stock all pay for the same reason: they carry equity-crash risk in a bond costume. Every one falls on the market’s worst days — preferreds fell 65% through 2008–09 — which disqualifies them from the one job the sleeve has.

−0.6% to −2.0% avg on QQQ’s worst days
On parole — awaiting forward evidence

DON55/20 trend following

forward-negative, capital-starved

Passed historical validation; its frozen forward paper stream is negative so far and skipped 11 of 13 signals for lack of capital.

re-heard when the account can hold 5 positions
On parole — awaiting forward evidence
A-RANKvalidated

Dealer gamma exposure (GEX)

a better forecast that loses money
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The parole hearing happened. With the options archive finally complete — every expiration, both symbols, the crisis years filled in — dealer gamma genuinely improves the volatility forecast, consistently, in both halves of the sample. Traded, it still loses to the plain 20-day rule: on the Nasdaq at identical turnover, so the loss cannot be blamed on costs. A forecast can be right on the days that don’t matter and wrong on the days that do.

passed the forecast gate · lost the trade in 19 of 19 configurations · both symbols
A-RANKvalidated

Implied-vol term structure

errors that breathe with the regime
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.
read the autopsy

The market’s own forecast of volatility — implied vol off the full chain — is the best predictor this project ever measured, six times stronger than dealer gamma. Fed into the sizing rule it lost anyway. Its errors run in one direction for months at a time, fattening and thinning with the very regime it is supposed to call, wrong precisely at the turns where the money is decided. The humble trailing estimate’s errors are pure delay, which the rule forgives cheaply; the option market’s are structural, which it cannot.

best forecaster ever tested here · still −17% of available headroom
B-RANKvalidated

The rest of the options surface (skew, vanna, charm, put/call OI)

same blood, same disease
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Crash-fear skew was empty outright — zero forecast content. Dealer vanna and charm forecast better than gamma and lost more when traded. Put/call positioning did nothing. Six option-surface signals, five passed a legitimate forecast gate, zero beat the control — across the family, the correlation between forecast quality and trading value came out negative.

the paid archive, fully mined: 0 for 6
A-RANKvalidated

Cross-market shock consensus

right, rare, and late
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

When equity options, vol-of-vol and credit all alarm at once, a shock usually follows — the one classifier whose precision survived every leave-one-year-out test we could throw at it. But it fires a day before the storm, catches roughly one event in seven, and slept through 2022’s grind entirely. Grant it hindsight-perfect responses to its exact alarms and the most it could ever add is a rounding error. The signal is real. There is no trade inside it.

35.6% precision, holdout-robust · perfect-action ceiling ≈ 0.3%/yr
A-RANKvalidated

Failed-re-entry memory (adaptive hysteresis)

a risk dial, not an edge
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

The strategy raising its own re-entry bar each time a re-lever fails — genuine adaptation, and it beat equivalent static caution in the 2000–02 grind it was aimed at. But everything it saves in long bears it hands back in sharp recoveries, and the Sharpe ratio came out identical to the second decimal. That is the signature of a preference, not an improvement — and the existing band already offers the same preference for free.

grind bears +9.8% · full 33 years −0.45%/yr · Sharpe unchanged
B-RANKvalidated

Levered-fund-only expression

beta in disguise
didn’t improve our engineTested as a change to our engine or an add-on beside it, judged on whether the engine did better with it than without it, at the same risk.

Express the whole position through the 2× fund at half weight and tracking error collapses — plus 1.2 points a year, says the naive backtest. Exposure-matched, it is 1.5 points worse: the extra return was nothing but more average exposure, and the fund’s daily-reset drag eats the tracking gain whole. The third idea this project has caught borrowing beta and calling it alpha.

+1.2% headline → −1.5% at matched exposure

Overnight-share recovery accelerator

promising — on forward parole

When overnight variance dominates inside an already-stressed tape, recoveries tend to follow — and accelerating re-entry on that signal beat both the control and a random-timing placebo in the modern era, 2022 included. But thirty years of history say the edge barely existed before 2019, and this graveyard holds too many era-local miracles to trust one more backtest. It now runs as a paper portfolio beside the live account — earning nothing, risking nothing, accumulating the only evidence that counts.

modern era +1.5%/yr · pre-2019 ≈ 0 · paper arm live in the forward ledger

Opportunistic insider buying

unprovable on free data — yet

Published α 0.72%/mo (t 2.27). Backtest blocked by survivorship; a forward collector has been accruing evidence nightly since Aug 2026.

verdict arrives in portfolio-years
The survivor
S-RANKlived
ARMED & LIVE — THREE MODES

The Distillate

vol-managed index core — what remained when everything else burned away

The one idea that passed every gate that killed the other one hundred and eighty-seven. Volatility is forecastable; direction is not; it acts only on the forecastable part. STEADY: no losing decade in 27 years. SELECT: the NASDAQ with the exit. ULTRA: the same, run hot.

16.2%
steady · modern era
23.6%
ultra · modern era
live record →

Methodology: honest fills (nothing fills at a price the signal made unreachable), day-clustered t-statistics, era splits on every claim, best-of-N permutation nulls over entire search spaces, matched placebo controls, and costs modeled at real account sizes. When our own tests were the bug — it happened four times in one night — that's on the stones too. ← back to the one that lived