Every desk that has been through an accounting blowup builds an accrual screen afterwards. Most of those screens are abandoned within eighteen months, and the reason is always the same. The screen fires on forty names a quarter, thirty five of them are fine, nobody has the analyst hours to check them, and the overlay quietly stops being run. The failure is not in the measures. It is in treating the output as a signal rather than as a triage queue with a known and measured error rate.
What follows is the version that survives, which means three measures that disagree with each other, a combination rule that trades sensitivity for workload, and an explicit protocol for measuring your own false positive rate rather than quoting somebody else's.
Three measures that fail differently
Run all three. The value is in the disagreement, because each one is defeated by a different manipulation and a company that trips two of them is a different proposition from one that trips one.
Total accruals scaled by average total assets is the blunt instrument. Net income minus cash flow from operations, divided by average assets over the period. High positive accruals mean reported earnings are running ahead of the cash the business generated. It is easy to compute across a full universe from standardised data, which is its main virtue, and it is noisy for any company with growing working capital, which is its main defect.
Receivable days drift against revenue growth is the sharpest of the three for revenue recognition problems specifically. Compute days sales outstanding each quarter, then compare its year-on-year change to revenue growth over the same window. Receivables growing meaningfully faster than revenue for two consecutive quarters is the pattern that precedes most channel stuffing and most aggressive percentage-of-completion recognition. It is also the pattern produced by a genuine shift into enterprise contracts with longer payment terms, which is why it needs the other two.
Cash conversion divergence is the slowest and the most reliable. Take cumulative net income over three years against cumulative cash flow from operations over the same three years. Over a single year the two legitimately diverge for a dozen boring reasons. Over three years, a company whose cumulative earnings exceed cumulative operating cash flow by a wide margin has either a structurally working-capital-hungry model or a recognition problem, and there is no third explanation.
What a single fundamental sub-score cannot be asked to do

None of this is a complaint about the sub-score. Compression is what a screen is for, and a fundamental score computed consistently across a universe that read 4,420 companies at capture, with 4,432 fully scanned, is doing something no desk replicates by hand. The point is architectural. Use the platform for universe definition and for the breadth pass, and run the accrual overlay as a separate layer on exported data where the three measures stay separate and auditable. The sector and index filters give you a stated population, the export gives you the population as a file, and everything after that is your own construction with your own version stamp on it.
The combination rule, which is a workload decision
Rank each measure into deciles within sector, never across the whole universe. Working capital intensity varies enormously by industry and a cross-sector accrual ranking is mostly a sector ranking.
Flag a name when it sits in the worst decile on two of the three measures in the same period. Requiring two rather than one is the difference between a queue you can work and a queue you abandon, and it is the single most consequential parameter in the design. Requiring all three makes the screen so specific it fires only after the problem is public.
Add one persistence condition. The flag has to survive two consecutive quarters before it enters the queue. A single quarter of bad accruals is frequently an acquisition, a contract timing shift, or a change in payment terms. Two consecutive quarters filters most of that at the cost of delaying the flag by a quarter, and that trade is worth taking on an overlay whose purpose is risk reduction rather than alpha capture.
Measuring your own false positive rate
The desk will want a number for how often the screen is wrong. Do not import one. Published error rates come from universes, periods and combination rules that are not yours, and the base rate of accounting problems varies enough across sectors and market regimes that a borrowed figure is worse than no figure, because it will be quoted in a memo as though it applied.
Measure it. The protocol is unglamorous and takes about a quarter to set up.
Define a true positive in advance and in writing. A workable definition is that within eighteen months of the flag, the company restated, or issued a material guidance cut attributed to receivables or revenue recognition, or changed auditor without an obvious structural reason, or wrote down receivables or inventory by a material fraction. Anything vaguer than that will be argued about retrospectively, always in the direction that makes the screen look better.
Then run the screen historically on point-in-time data. This is the step that is skipped and it invalidates everything when it is. Restated financials are the single largest source of look-ahead bias in accounting research, because the restated numbers show the problem the screen was supposed to detect. If your data vendor cannot give you the figures as originally filed, your backtested hit rate is not a hit rate.
Report the confusion matrix quarterly with the base rate alongside it. The base rate is what makes the number interpretable. If genuine accounting events occur in a low single digit percentage of names per year, then a screen with excellent statistical properties will still produce a queue that is mostly false positives, and this is arithmetic rather than a fault in the screen. A desk that understands that keeps the overlay running. A desk that expected the screen to be right most of the time cancels it.
What a flag entitles the desk to do
Decide this before the first flag, because deciding it while looking at a specific position is how the overlay becomes decorative.
A flag is not a sell instruction and it is emphatically not a short mandate. What it should trigger is a position cap at some fraction of the normal size, a mandatory analyst review within a defined window with a written conclusion, and an escalation to the risk committee if the name is above a size threshold. The written conclusion is the part that compounds, because after two years you have a file of reviewed flags with outcomes, which is what turns the false positive rate from an argument into a measurement.
Where the overlay systematically misfires is worth stating in the same document. Fast-growing companies with genuinely expanding working capital will flag repeatedly and correctly on the mechanics while being fine on the substance. Serial acquirers flag because consolidation moves working capital discontinuously. Businesses that finance their customers flag structurally. Banks and insurers should be excluded outright rather than scored badly. Write the exclusions down as policy, because an exclusion applied at analyst discretion when a favoured name flags is the exact failure the overlay exists to prevent.