Every cluster hit rate I have been handed by a manager was computed the same way. Pull the cluster events, join them to a price series, compute forward returns, count the winners. The join is where the number dies, because the price series almost always comes from a vendor table keyed on currently listed tickers, and the names that stopped being currently listed are exactly the ones carrying the information. What comes out the other side is a statistic about companies that survived, presented as a statistic about companies you would have bought.
This matters more for insider clusters than for most signal families. Clusters concentrate in small companies with concentrated registers, and small companies with concentrated registers are the population that gets acquired and the population that runs out of money. Both outcomes remove the ticker. Neither outcome is a neutral exclusion.
The universe a live feed can see
The coverage statement on the module is present tense and accurate about itself. Every SEC Form 4 filing from US listed companies, officer, director and ten percent owner transactions, streaming, with a sixty second refresh and a US Form 4 scope. That is a description of a live feed, and a live feed is by construction a live universe. The filings of a company that was taken out three years ago have not been deleted from the record, but they are not what a streaming panel is built to show you, and your backtest inherits whatever universe your extract was drawn from.
So the numerator is easy and the denominator is the whole job. At capture the header strip showed 312 multi-insider clusters over the trailing seven days. Signals at that rate are not scarce. What is scarce is a defensible list of every name that was listed on each historical date, which is the object you actually need and the object nobody hands you.

Note one more thing about that strip before you lift figures off it into a research note. The buy and sell counters read 176 and 788 filings over seven days while the cluster counter reads 312 and the discretionary counter reads 2,824. Those tiles are not computed over a shared denominator, so ratios built across them are meaningless. That is fine for a dashboard, which is doing a different job. It is not fine as an input, and the general rule it implies is worth internalising. Recompute every aggregate you intend to publish from the underlying records rather than from a panel that was designed to be glanced at.
Two errors that point in opposite directions and do not cancel
The first is truncation at acquisition. A cluster fires, the company is acquired eight months later at a premium, and the price series ends on the delisting date. Depending on how your pipeline handles that, the observation is dropped, or carried forward flat, or filled with nulls that get silently excluded by a mean. In every one of those cases you have deleted an outcome that was, for a strategy built on insiders buying their own shares, plausibly among the best outcomes in the sample. Insider clusters preceding a takeout are not a rare pathology. They are one of the mechanisms the strategy is supposed to be capturing.
The second is truncation at failure. A cluster fires, the company dilutes, the shares get delisted to an over the counter venue, and the final print is thin, wide and unreliable. That observation is often dropped too, and it is dropped for exactly the reason it should be kept. A minus 90 percent outcome removed as bad data flatters your base rate directly.
Practitioners often assume these two effects roughly offset. They do not, because there is no reason for the counts to match and no reason for the magnitudes to match, and the balance shifts with the cycle. In a heavy takeout year the survivorship error runs one way and in a funding drought it runs the other. The only defensible position is to handle each terminal state explicitly and report how many observations each one contributed.
A terminal state for every cluster
The construction I would defend in a review is boring. Every cluster event gets a row, every row gets a terminal state, and every terminal state has a stated return rule that was fixed before anyone looked at the outcome.
| Terminal state | Return rule | What it protects against |
|---|---|---|
| Still listed at horizon | Close to close over the window | Nothing, this is the easy case |
| Acquired for cash | Deal price on the effective date, then hold in cash to horizon | Deleting your best outcomes |
| Acquired for stock | Roll into the acquirer at the exchange ratio | Fake truncation on a position you would still hold |
| Delisted to an over the counter venue | Last reliable print, then the OTC series if you would have held it | Silent removal of bad outcomes |
| Liquidated or cancelled | Minus 100 percent | Optimism by omission |
Key the whole thing on the filer and issuer identifiers rather than on tickers. Tickers get recycled, reassigned after bankruptcy and reused by unrelated companies, and a ticker keyed join will quietly attach one company's cluster to another company's price history. Form 4 filings identify their filers and issuers by SEC identifier, so the clean key is available to you. Use it.
What the performance view can settle and what it cannot
The module's forward return attribution view is the natural place to look for a house base rate, and it is worth reading carefully rather than trustingly. At capture, with the filters at their default breadth, it reported 20,000 scored filings, a seven day forward return of plus 0.25 percent against plus 0.07 percent for SPY, and a 46 percent win rate. The 30 day, 90 day and 365 day panels each reported zero scored filings.
Three readings follow. First, a 46 percent win rate on the broad filing set is a useful corrective to anyone who thinks the raw feed is a strategy. It is not, and the panel is not pretending otherwise. Second, and more relevant here, the horizons where survivorship actually bites are the long ones, because a takeout or a liquidation takes quarters rather than days. Those were the horizons showing no scored filings on the day I looked, which means that panel could not have answered the survivorship question even if it had been designed to. Third, the same header reported a notional figure of 1,089.83 trillion dollars alongside those 20,000 filings, which cannot be a true sum. Aggregates on a monitoring screen sometimes carry unit or scaling artifacts. Reconcile before you cite.
None of that is a reason to avoid the module. It is a reason to be specific about which questions a live panel is built to answer. Monitoring and measurement are different products, and the failure mode on desks is using the first as though it were the second because it is already on screen.
The version of the number that survives a review
When the base rate goes into a memo, it goes in with five things attached, and a number missing any of them should not clear internal review. The event count. The exact window and horizon. The terminal state breakdown, including how many names were acquired and how many were delisted for cause. A confidence interval, because a hit rate on a few hundred events has a range around it wide enough to change the decision. And the exclusion log, meaning every filter you applied and how many events each one removed.
That last item does most of the work. In practice the difference between a survivorship contaminated base rate and a clean one is not sophistication, it is whether anyone wrote down what got dropped. A hit rate quoted without an event count and without a delisting count is not a base rate. It is a marketing number that has learned to dress like one, and it will be quoted back at you in the meeting where the strategy is being defended.