Detecting Earnings Manipulation With the M-Score

The eight variables in the M-Score model are designed to capture bloating or cosmetic changes made to a company’s financial statement numbers.

Messod D. “Daniel” Beneish, Ph.D., a professor of accounting at Indiana University, created a screening model that helps investors, lenders and auditors detect companies most likely to have misleading financial reports. His M-Score (M for manipulation) has served as an early warning bell for notorious manipulators both domestically and internationally (Enron in the U.S., Satyam in India, Wirecard in Germany and Kangmei Pharmaceutical in China). We spoke about how the M-Score works and, most importantly, how individual investors can use it to identify quality growth companies for their portfolios.
—Cynthia McLaughlin

What is the M-Score and what is it designed to do?

The M-Score is a screening model. It is designed to detect manipulation in financial statements. It originated in examining cases of firms that violated accounting rules and reported numbers that were misleading. The U.S. Securities Exchange Commission (SEC) charged the majority of these firms with accounting fraud. Other firms were identified because the results of some internal investigations were made public. These internal investigations were generally initiated by new members of the board of directors who ended up discovering irregularities and misstatements in financial statements.

What are the components of the M-Score and how is it calculated?

I estimate the probability that a firm is manipulating its financial statements by creating a model to compare two different types of firms: manipulators and fraud firms, identified with a score of 1, and other firms in the industry, identified with a score of 0. For all practical purposes, the model includes variables that either captured the bloating in financial statement numbers that arise as a consequence of the manipulation or revealed the incentives that could have prompted a firm to make cosmetic changes to its books. That’s the spirit of what I was trying to capture with my model variables.

How many variables are there in the calculation?

There are eight variables. One variable that I look at to assess ‘bloating’ is based on receivables. The model views receivables growing at a rate much faster than sales as an indication that there are potentially some revenues that have been recognized too early or fictitiously. In either case, those revenues are not going to be collected or, at least, not anytime soon.

When there is an imbalance in how receivables are growing relative to sales, that is suggestive of revenue manipulation. This is an example of something we could observe. I think managers who are exposed to these dire situations have become smarter about this. They may, for example, factor receivables, which is the process of selling accounts receivables to banks at a discount. Factoring makes it possible for a manager to manipulate revenues and to decrease receivables by selling some of their viable receivables. Of course, most analysts, fraud examiners and other informed financial statement readers would know to look in the footnotes to observe changes in this sort of behavior. That was an example of bloating.

An example of an incentive variable is relying on the expectation that gross margins [revenue less cost of goods sold] shouldn’t change unless there’s a fundamental change in the business model. If I expect gross margins to stay the same and see that gross margin this year is lower than last year’s, I think that this deterioration impacts the bottom line. In turn, that provides firms with greater incentives to make adjustments that increase income. Those are two examples of variables in the screening model.

The M-Score Calculation

Eight Variables

  • Days’ Sales in Receivables Index (DSRI):
    (Accounts Receivable Y1 ÷ Sales Y1) ÷ (Accounts Receivable Y2 ÷ Sales Y2)
  • Gross Margin Index (GMI):
    (Gross Income Y2 ÷ Sales Y2) ÷ (Gross Income Y1 ÷ Sales Y1)
  • Asset Quality Index (AQI):
    [1 – (Current Assets Y1 + Prop, Plant & Equip Y1) ÷ Total Assets Y1] ÷ [1 – (Current Assets Y2 + Prop, Plant & Equip Y2) ÷ Total Assets Y2]
  • Sales Growth Index (SGI):
    Sales Y1 ÷ Sales Y2
  • Depreciation (DEPI):
    [Depreciation Y2 ÷ (Depreciation Y2 + Prop, Plant & Equip Y2)] ÷ [Depreciation Y1 ÷ (Depreciation Y1 + Prop, Plant & Equip Y1)]
  • Sales, General and Administrative Expenses (SGAI):
    (SGA Y1 ÷ Sales Y1) ÷ (SGA Y2 ÷ Sales Y2)
  • Leverage Index (LVGI):
    [(Long-Term Debt Y1 + CL Y1) ÷ Total Assets Y1] ÷ [(Long-Term Debt Y2 + CL Y2) ÷ Total Assets Y2]
  • Total Accruals to Total Assets (TATA):
    (Income Before Unusual Items and Discontinued Ops. Y1 – Cash Flow From Operating Activities Y1) ÷ Total Assets Y1

M-Score Calculation

= –4.84 + (0.92 × DSRI) + (0.528 × GMI) + (0.404 × AQI) + (0.892 × SGI) + (0.115 × DEPI) – (0.172 × SGAI) + (4.679 × TATA) – (0.327 × LVGI)

Y1 = current year; Y2 = prior year.

Source: “Earnings Manipulation and Expected Returns,” by Messod D. Beneish, Charles M.C. Lee and D. Craig Nichols; Financial Analysts Journal, Volume 69, Issue 2 (2013).

What inspired you to create the M-Score?

My teaching was the inspiration. I was teaching at Duke University and found it difficult to actively engage students on accounting topics in a financial statement analysis class for students pursuing MBA degrees. When I once used an article from The Wall Street Journal that was critical of the accounting practices at a particular firm, they sort of woke up. All of a sudden, they were interested in understanding the accounting behind what The Wall Street Journal was saying—which often was simply a reflection of what the SEC had said in court records, or what had been revealed through an internal investigation. So, I incorporated a few more articles discussing accounting issues and found that students were interested and receptive to learning about different accounting methods as a result.

I thought this was interesting and wondered if there were more. I did a wide search and got a small number of hits from the Dow Jones News Service. Of course, some of them were articles that discussed nothing specific about manipulation: They were articles written by a short seller who likely wanted to entice investors to sell the stock. For this reason, in estimating the model I focused on public admissions of the need to restate to correct some misreporting or on links to SEC cases. What started as a teaching project became a research exercise. That is how it started.

How long did that research take to complete?

It took nearly three years as I presented the paper at many places, received feedback and made revisions. My intent was to convince researchers that this was a valid prediction model.

I was given the runaround at two of the top journals. Actually, one top journal rejected my paper after five rounds, which is rather unusual. I think the editors rejected it because they thought the model worked on a small sample of firms, and it wasn’t going to work afterward. I could not have done anything else to show that the model would work in the future, so I didn’t get play at the very top journals.

But then, two influential academics called me because they had seen an article in The Wall Street Journal about a firm that had committed fraud, had run the numbers through the model and told me, “Well, this works.” The model performance in real time convinced several academics much more than anything I could have said. These became the advocates for the paper and they started teaching the model and incorporating it in textbooks. The rest is history. Since 1999, as we have shown in a recent article, there hasn’t been a more economically viable model for investors.

Is there a target M-Score that is considered optimal? Or should we think about an M-Score that is greater than or less than the negative 1.78 value?

If a firm’s M-Score is greater than –1.78, I treat it as a likely manipulator. That is how I like to think about it, that is how I use it. Since the model provides a number that fits under normal distribution, –1.78 corresponds to a normal probability of 3.75%. It is a low probability of manipulation. But at the time I was doing this research, the probability of manipulation in the economy in the sample I had was 0.69 of 1%. A probability of 3.75% was more than five times greater than the average firm in that sample.

While I rely on probabilities that are coming out of the model, I also rely on what each variable tells me. I’ve worked with some investment firms and some accounting firms and the advice I give is to look at the neutral values of the variables in the model. If you see values greater than 1, there is a potential problem, either with bloating of assets or with incentives to inflate earnings. Each variable essentially tells you where to look.

One of the variables I have, for example, is the rate of depreciation last year compared to this year. If a company depreciated more aggressively last year compared to this year, and the ratio is greater than one, what does that tell me? That tells me that this year, the firm did what it could to reduce its depreciation expense. Google has taken some of these measures recently by increasing the service life of some of its assets. This is not illegal by any means because it’s disclosed. People who read the footnotes will know that the life of service is now expected to be longer and how this affects net income. Of course, not everybody reads the footnotes; they are more costly to access and read since there is no standard format. Although, with web crawlers and artificial intelligence (AI) tools, it is getting less costly to access footnote information.

Some of our members are interested in growth investing, and we have a model portfolio focused on growth companies. Can investors use the M-Score to identify quality growth companies?

Sales growth is a variable in the model. The idea is that when a firm is growing fast, it has an interest in maintaining a pattern of revenue growth. Usually, when there is a deceleration, the stock price tanks. So, sales growth is in the model because high-sales-growth firms have this incentive. Management in high-sales-growth firms also have greater ability to manipulate. This is because in high-growth firms, retaining first-mover advantage often means that control concerns are not as important as production to meet demand.

Generally speaking, the firms flagged by the M-Score tend to face some headwind. The false positives identify the non-fraud firms, and those identified tend to have poor returns one year after. Yet, because of the sales growth variable, the M-Score will flag growth firms proportionately more than other firms. The question is, how do you tell them apart? Well, one way is to look at governance and monitoring. Is there a credible board? What do auditors say about the existence of internal control or material control weaknesses? Another way is to look at what is happening with variables used to calculate the M-Score. For example, the gross margin index, or accruals to assets or the accounts receivable to sales variable is what I would look at.

Is there a correlation between firms that have high M-Scores and asset bubbles? There has been recent talk about stock market asset bubbles.

When we were discussing research at conferences and meetings in the late 1990s, I remember thinking that there were no earnings behind any of the internet firms. Some of my colleagues argued this was not unusual as that’s how an industry begins. I remember thinking these metrics were earnings before expenses because if you tried to look at anything beyond sales, you got lost in negative numbers. We know what happened with the bubble of the early 2000s. I was right, and I was wrong. I was right to see that there was some overvaluation. I was wrong to think that this would be the end of it, as the internet is here to stay.

The 2008 crash was primarily driven by banks and, as the M-Score does not apply to banks, I do not have much to say here. The balance sheet structure of banks is so different that none of the fraud models I have seen apply to banks or insurance companies, or real estate investment trusts (REITs).

How do you see individual investors using the M-Score for their strategies and portfolios?

If you apply the M-Score to 2,000 firms right now, you are more likely than not to see between 120 and 150 firms flagged [score higher than –1.78]. That is, if you’re filtering through 2,000 firms, you are going to have between 6.0% and 7.5% flagged. Of these, between 10 and 15 firms are actual frauds while the remaining 100 to 120 firms are typically facing some headwind.

I would say at first pass that avoiding those firms is a good idea. However, if some of the flagged firms are those that an investor is interested in, then I would look at the details of the M-Score and do a more careful analysis. By that I mean, at the very least, looking at trends on multiples, looking at a decomposition of return on assets and comparing the firm to itself in prior years and to a competitor in the current year.

Note that there are other models that would do a better job than the M-Score at identifying fraud firms, perhaps identifying between 20 to 25 frauds. A problem is that those models will also flag 775 out of 2,000 firms as fraudulent that do not misreport. If you avoid these 800 firms, what my prior work shows you is that you are going to miss out on a lot of price appreciation. In addition, the analysis is costly because you don’t know who the 25 frauds are within the 800. That is, the number of frauds identified is small relative to all the firms that are flagged, which makes them costly to identify. You would drown just trying to find the frauds. That’s a difficulty.

Indeed, the superior economic viability of the M-Score isn’t because it has the highest hit rate [success rate at identifying fraud], it is because my model has the lowest false flag or false positive rate, and the false positives that look like frauds are facing headwinds.

Given that the M-Score can produce a large field of investments to consider, are there other scores that can be utilized along with the M-Score that would help whittle this down, such as the F-Score?

The F-Score has a 38% false positive rate. Rounded down to 35% on 2,000 firms, that is 700 firms with a flag. That is not going to help. I have not evaluated research models that use financial statement data along with textual analysis. Those don’t have hard and fast rules. If you took the firms identified by the M-Score and did some text analysis on them just using ChatGPT, and asked, “Do you see any problems here?” and read the conclusions, this could help, but I have not done this.

The models that have done both text analysis and number analysis, like the M-Score or the F-Score, have a 75% success rate in identifying fraud. That’s huge. That means they identify three out of four fraud cases. Now, how many cases of fraud firms are there in the past 40 years? About 400, so that 75% is equivalent to 300 discovered frauds. The cost is that those models are even worse at generating false positives, with false positive rates as high as 58%. Imagine running a medical test where you have 58% false positives. Nobody has the resources to do that follow-up. Hopefully, my colleagues will start following my suggestion to develop models with the goal of reducing the number of false positives. I have seen a new method that uses k-nearest neighbor technology [machine-learning technology], like a Google technology. When you do a search on Google, what is the first thing that shows up? This has promise in terms of reducing the number of false positives.

Your 1999 study summarizes that the model can’t reliably be used to study companies operating under circumstances that are conducive to decreasing earnings. Can the M-Score be used to detect big bath accounting [taking one-time charges to lower current net income] or predict similar situations?

No, because it was estimated exclusively on what I observed as enforcement back then, which were earnings overstatements. It cannot be used for that application. It can be used for private firms. It can be used for foreign firms and has been used worldwide, but it cannot be used for income reduction or understatement. The only thing I would venture to say is that a firm that takes a big bath is oftentimes a firm that inflated its assets and earnings in prior years.

Your recent research examines the power of misreporting and measuring if the U.S. is going to have a recession. Regulators and forecasters can get immediate use out of this application. But are there takeaways for individual investors?

None, beyond the rationale that links what we call the aggregate M-Score, which is a value-weighted M-Score of approximately 2,000 companies that announce within 90 days of the end of a calendar quarter. Banks are not included, obviously, because we couldn’t compute the score for banks. We call this measure the likelihood of misreporting in the economy.

The link is that accounting information is not only useful for investors, but it’s also useful for making investment decisions, production decisions and hiring decisions in product and labor markets. What we’re saying is that the financial information of certain firms is faulty, and the more firms that have faulty information in the economy, the more their competitors who are not misreporting will make mistakes such as overinvesting or overhiring. This is because it is not yet known that the financial information of some firms is faulty, so firms only recognize with a delay that the economy is slowing. As things unravel, peers of manipulators recognize that some bad times are coming and they start cutting investments and reducing their workforce. This basically reduces household consumption and the level of investment.

The other consideration is that we have a probability of between 60% and 70%, depending on what quarter we’re measuring, that there will be a recession, either in quarter four of 2023 or anytime in 2024. There is no certainty, but the highest probability of a recession we estimated was 69% last quarter, and this quarter it is slightly down to 61%, but still fairly high. 

Discussion

MARK H from CAN posted over 2 years ago:

Sounds like an excellent tool. Please provide a website that uses it!!!


MESSOD B from IL posted over 2 years ago:

My M-Score calculator is pretty friendly and can be found at: https://apps.kelley.iu.edu/Beneish/MScore/MScoreInput


F K from SC posted over 2 years ago:

Thanks.


BARRY J from TX posted over 2 years ago:

Cynthia The one question I wished you had asked Dr. Beneish is how many times have you testified in a court under oath about findings from your model and what was the outcome of the case? That data would make me feel a lot more confident about the outcome of a model complex based on probabilities calculated from the interaction of 8 variables that are (by definition) highly subject to manipulation. Probabilities possibly indicate a correlation, but do not necessarily indicate causality. The standard of proof in a civil court is a low bar to pass but that data would provide more certainty as to how its application serves to predict intent and culpability. "It ain't a bird dog until it brings home a dead bird."


BARRY J from TX posted over 2 years ago:

I wonder how many "hits" we would get (M < -1.78) if we applied the M Score procedure to the database used in the first step of the AAII stock screening process. The issue of "false positives" (mentioned in the article) is also a concern. If not all AAII screens, the Growth Income Platinum offering seems to be where potential suspects may be lurking in the early screening steps. Maybe Dr. Beneish can help us do this. I only ask for a number of hits, not any names, and would like to see the amount of time it takes to perform these screens. Although potentially valuable, this complex process seems time-consuming and cumbersome for anyone who is not a distinguished professor of accounting who may have access to more research capabilities than the average AAII member.


Kyle I from CAN posted over 2 years ago:

I have found another website that utilizes the Beneish M-Score: https://www.gurufocus.com/term/mscore/AAPL/Beneish-M-Score/Apple


MESSOD B from IL posted over 2 years ago:

In addition to Bloomberg. Capital IQ and Audit Analytics, below are a few other sites that either speak of or use the M-Score: Wikipedia - https://en.wikipedia.org/wiki/Beneish_M-score Investopedia - https://www.investopedia.com/terms/b/beneishmodel.asp Seeking alpha - https://seekingalpha.com/article/3293805-the-m-score-explained Stockopedia - https://www.stockopedia.com/ratios/beneish-m-score-ttm-5360/ Corporate Finance Institute - https://corporatefinanceinstitute.com/resources/financial-modeling/beneish-m-score-calculator/ Refinitiv - https://developers.refinitiv.com/en/article-catalog/article/Beneish-M-Score-and-Altman-Z-Score-for-analyzing-stock-returns-of-the-companies-listed-in-the-SP500 Wall Street Mojo - https://www.wallstreetmojo.com/beneish-m-score/ Quant Investing - https://www.quant-investing.com/glossary/m-score-beneish Business Insider - https://www.quant-investing.com/glossary/m-score-beneish Day Trading - https://www.daytrading.com/beneish-m-score


Wayne T from IL posted over 2 years ago:

@Barry J Using the GuruFocus M-Score calculator, all but one of the 20 stocks in the AAII Growth Investing model portfolio had an M-Score of -2.21 or lower as of September 7. According to Dr. Beneish, an M-Score less than or equal to -1.78 suggests that the company is unlikely to be a manipulator. This data serves to support the Growth Investing strategy of identifying quality, sustainable growth companies. I hope this information is helpful. Wayne A. Thorp Creator, AAII Growth Investing


BARRY J from TX posted over 2 years ago:

Thanks, Wayne. I just knew you were itching to run the numbers on the AAII Growth data. I agree with your conclusion. Although we have no SDs to standardize the magnitude of any differences, these differences look "significant" at face value.


MATTHEW C from UT posted over 2 years ago:

Great article, but could the author please provide the M score for a set of companies (like Barry J in comment above) in order for readers to learn something regarding stocks in the current market? I think most readers are not going to be inclined to manually input all these parameters for themselves for more than a handful of stocks. Thanks!


ROBERT A from NC posted over 2 years ago:

Note to AAII: Adding the M-Score as a factor on SI Pro would increase its value!


RONALDO J from IL posted over 2 years ago:

Thanks very much for the interview and article. Just a few comments. Comparing the F-Score and M-Score may be inappropriate since they were developed to solve different problems. One very important use for M-Score is to screen out what Warren Buffett calls "Too Hard" investment prospects. If you get a M-Score flag and you cannot figure out why a company is doing something with its books then stay away. Further, I would recommend that every quarter AAII publish the list of M-score flagged companies as a public service (If we can help investors avoid an Enron situation everyone would benefit.)


WILLIAM C from TX posted over 2 years ago:

I agree with Robert A. Why not put the M-Score in SI Pro?


DAVID G from CA posted over 2 years ago:

Thanks for this insightful and practical article. However, I have mathematical quibble. In the discussion on probability of deviation from the mean, it asserts that 3.75% is five times greater than .69%. The larger number is in fact five times as large as or, equivalently, only four times greater. But who’s counting?


STEPHEN H from FL posted over 2 years ago:

I agree with Robert A & William C. I built an M-Score custom field for SI Pro but it would be nice to have it as part of the stock package.


You need to log in as a registered AAII user before commenting.
Create an account

Log In

Get your free copy of our special report analyzing the tech stocks most likely to outperform the market.

Download the FREE Report Here: