In plain words: A review of stock-prediction projects pinpoints five recurring errors — too little data, wrong scaling, mishandled time order, vague targets, and misleading scores — and shows how each skews results. Correcting these steps is the way to make such market models trustworthy for real trading firms.
Abstract · Common Mistakes when Applying Computational Intelligence and Machine Learning to Stock Market modelling
For a number of reasons, computational intelligence and machine learning methods have been largely dismissed by the professional community. The reasons for this are numerous and varied, but inevitably amongst the reasons given is that the systems designed often do not perform as expected by their designers. The reasons for this lack of performance is a direct result of mistakes that are commonly seen in market-prediction systems. This paper examines some of the more common mistakes, namely dataset insufficiency; inappropriate scaling; time-series tracking; inappropriate target quantification and inappropriate measures of performance. The rationale that leads to each of these mistakes is examined, as well as the nature of the errors they introduce to the analysis / design. Alternative ways of performing each task are also recommended in order to avoid perpetuating these mistakes, and hopefully to aid in clearing the way for the use of these powerful techniques in industry.
E. Hurwitz, T. Marwala
arXiv:1208.4429 · stat.AP, cs.CY, q-fin.GN · submitted Aug 22, 2012
abstract · pdf · 5 pages
- Most ML algorithms perform well when there is some underlying phenomenon whose characteristics are being statistically inferred by the model. However, the stock market is now essentially a large network of computers trying to model one another. You can imagine how this might break some of the underlying assumptions that good ML results rely upon. We can see the results of this in the increasing frequency of "flash crashes" caused by over-leveraged quant hedge funds all tripping over sell triggers and/or getting margin calls at the same time.
- In practice, asset class/trading strategy correlations tend to change dramatically over time. Nowadays we talk about the market being 'risk-on' or 'risk-off' modes, since we're used to seeing market-wide selloffs and buyins. Your ML model can only do so much if the entire market is selling everything, unless you're going to dip your toes in derivatives or short selling to which I say: good luck! :)
- A major, major issue is execution. You can have the greatest and most accurate model in the world, but actually trading it is another beast entirely. Bid/ask spreads and market movements from your trading, particularly if you are dealing with a non-trivial amount of money or trading securities that have less than ideal liquidity, is usually going to eat up any alpha you might have. In my own backtests of fairly straightforward trading algorithms even a minor 0.05% spread on bid-ask spreads for a weekly trading algorithm can eat your lunch, nevermind if you're planning on doing intra-day trading or trading anything other than the most popular funds/stocks.
- Beyond any of these risks, you're going to have to inevitably suffer through downdrafts. I don't know about you, but watching my money disappear as it's being controlled by a trading algorithm/model that is subject to all kinds of mistakes and bugs is well beyond my own intestinal fortitude.