Research meetings on automated strategies tend to circle one unstated confusion. Somebody reports that the profile improved, and nobody in the room can say whether the entry logic got better or the risk wrapper around it got luckier. Those are different accomplishments with different half lives, and the second one decays much faster than people expect, because a wrapper tuned on a window is a wrapper fitted to that window.
The confusion is structural rather than careless. A backtest returns one number, and that number is the composition of a signal, a sizing rule and a stack of constraints, applied in an order that matters. Untangling it takes a deliberate design and about a dozen runs.
On a follower profile, the desk contributed no signal at all
Start with a distinction that the profile list makes unusually clear. On this desk every enabled profile is a follower, named for the source it tracks: an index profile following a momentum engine on a daily cadence, single stock profiles doing the same, a silver pair and a crypto pair on four hour cadences, plus one reversal engine test. Eleven profiles, ten enabled.
Read that literally and the attribution question becomes sharper than it usually is. None of these profiles generate their own entries. The signal is produced elsewhere and consumed here, which means every parameter the desk has touched, the sizing rule, the filters, the leverage cap, the risk guards, the per source rules, sits entirely on the constraint side of the ledger. When the profile's result changes and the source did not, the change is attributable to the cage by construction.
That should make the analysis easier and in practice it makes the temptation worse, because the only thing the desk can adjust is the cage, and a team that can adjust one thing will adjust it until the number improves. Naming that dynamic is the first control. Measuring it is the second.

The ladder of nested runs
The Backtest a Profile panel takes a profile, a number of days and a start equity, and replays historical signals with no real orders, honoring all filters, sizing, leverage caps and risk guards. There is no attribution view in the panel and I would not assume one exists, so the decomposition is built by running a ladder of profile copies and differencing the results.
Construct the ladder as nested versions of the same profile, each adding one layer.
- Signal only. Every filter off, guards set wide enough not to bind, a fixed unit size on every entry. This run is your estimate of what the source produced before you touched it.
- Signal plus universe filters, so instrument and condition screens are active but nothing else.
- Signal plus filters plus the sizing rule, whatever converts a signal into a position size.
- Add the leverage cap.
- Add the risk guards one at a time, in a fixed order you write down.
- The production profile, with everything on. This should reproduce your existing result, and if it does not, stop, because one of your copies differs from production in a way you did not intend.
Run every rung on the same day with the same day count and the same start equity. The day count is measured backward from the moment you press the button, so a ladder built over three days is a ladder with a hidden variable in it. Record trade count alongside result at every rung, because a layer that removes a third of the trades has changed the statistical weight of everything downstream and a result reported without its count is not comparable to the one above it.
The differences between consecutive rungs are your attribution. The first rung is signal. Everything after it is constraint, itemised.
The ordering problem, and the convention that fixes it
The ladder has a known weakness that you should raise before a reviewer does. The layers do not commute. A concurrency limit applied before a volatility filter removes a different set of trades than the same limit applied after, because each layer changes which signals are still live when the next one evaluates. So the attribution you report depends on the order in which you added things, and any order you pick is a convention rather than a fact.
There are two defensible responses and one indefensible one, which is to pick the order that makes your favourite layer look best and not mention that a choice was made.
The first defensible response is to fix the order once, in a document, and use it for every ladder the desk runs. Order by how upstream the layer is: universe, then sizing, then hard caps, then discretionary guards. Attribution is then comparable across profiles and quarters, which is most of what you wanted from it.
The second is to run several orderings and average each layer's marginal contribution across them, which removes the dependence on any single sequence at the cost of a lot more runs. On a six layer profile the full set is impractical by hand, but four or five sampled orderings will tell you whether the ordering effect is material. If the spread across orderings is small, report the fixed order result and note the spread. If it is large, that is itself the finding: your layers are interacting strongly, and any single attribution number is misleading no matter how it was produced.
What the split tells the desk to do next
The decomposition earns its keep in the decisions it forecloses.
If the signal only rung carries almost all of the result and the constraint layers subtract modestly, the desk's job is capacity and cost, not tuning. The wrapper is working as intended, the marginal return to further guard adjustment is low, and effort should go to execution quality, instrument selection, or negotiating with the source about coverage.
If the signal only rung is weak and the constrained profile looks respectable, you have a problem rather than an achievement. A constraint stack that turns a poor signal into a good result has almost certainly been fitted, because constraints have no independent alpha. They can only remove trades. A stack that removes exactly the losing ones is a stack chosen after seeing which ones lost. The test is a clean out of sample window with the stack frozen, and the expectation should be that most of the improvement does not survive it.
If one guard accounts for most of the constraint effect, isolate it and check how many times it bound. A layer contributing a large share of the result while firing nine times over the window is an estimate built on nine observations, and the correct report is the contribution alongside the firing count, so the reader can weigh it themselves. This is the single most common way a research note overstates a risk control.
The layers the replay does not price
Two components of the live result are absent from every rung of the ladder, and both sit on the constraint side, which means an attribution built only from replays systematically understates what the cage costs.
The first is execution. The replay places no real orders. The module states directly that with paper mode off it places real orders on connected exchanges, that execution prices may differ from signal prices because of market conditions and latency, and that the user is solely responsible for the trades, with paper mode first and small live sizes as the prescribed route. Constraints interact with execution in a specific and unhelpful way: a cap that binds during a cluster does not decline a random trade, it declines the trade you were least likely to fill well anyway, so the measured cost of the guard and the measured cost of slippage are not independent.
The second is intervention. Every manual pause, resume or use of the emergency kill switch is a constraint applied by a person, and it appears in no profile configuration and no replay. If the desk has halted a profile three times this year, those halts belong in the constraint column with the others, priced by the same counterfactual method: what the profile would have done had it kept running. That number is frequently the largest single line in the decomposition, and it is the one nobody writes down. The failure mode is not that discretionary halts are wrong. It is that a desk which does not measure them will keep believing it is running an automated strategy when it is running a supervised one with an unrecorded overlay.