In the search for uncorrelated strategies and alpha, fund managers are increasingly adopting quantitative strategies. Beyond strategies based on alternative risk premia, a new source of competitive advantage is emerging with the availability of alternative data sources as well as the application of new quantitative techniques of Machine Learning to analyze these data. This ‘industrial revolution of data’ seeks to provide alpha through informational advantage and the ability to uncover new uncorrelated signals. The Big Data informational advantage comes from datasets created on the back of new technologies such as mobile phones, satellites, social media, etc. The informational advantage of Big Data is not related to expert and industry networks, access to corporate management, etc., but rather the ability to collect large quantities of data and analyze them in real time. In that respect, Big Data has the ability to profoundly change the investment landscape and further shift investment industry trends from a discretionary to quantitative investment style.
There have been three trends that enabled the Big Data revolution:
- Exponential increase in amount of data available
- Increase in computing power and data storage capacity, at reduced cost
- Advancement in Machine Learning methods to analyze complex datasets
Exponential increase in amount of data
With the amount of published information and collected data rising exponentially, it is now estimated that 90% of the data in the world today has been created in the past two years alone. This data flood is expected to increase the accumulated digital universe of data from 4.4 zettabytes (or trillion gigabytes) to 44 zettabytes. The Internet of Things (IoT) phenomenon driven by the embedding of networked sensors into home appliances, collection of data through sensors embedded in smart phones (with ~1.4 billion units shipped in 2016), and reduction of cost in satellite technologies lends support to further acceleration in the collection of large, new alternative data sources.
Increases in computing power and storage capacity: The benefit of parallel/distributed computing and increased storage capacity has been made available through remote, shared access to these resources. This development is also referred to as Cloud Computing. It is now estimated that by 2020, over one-third of all data will either live in or pass through the cloud.
A single web search on Google is said to be answered through coordination across ~1000 computers. Open source frameworks for distributed cluster computing (i.e. splitting a complex task across multiple machines and aggregating the results) such as Apache Spark have become more popular, even as technology vendors provide remote access classified into Software-as-a-service (SaaS), Platform-as-a-service (PaaS) or Infrastructure-as-a-service (IaaS) categories. Such shared access of remotely placed resources has dramatically diminished the barriers to entry for accomplishing large-scale data processing and analytics, thus opening up big/alternative data based strategies to a wide group of both fundamental and quantitative investors.
Machine Learning methods to analyze large and complex datasets
There have been significant developments in the field of pattern recognition and function approximation (uncovering relationship between variables). These analytical methods are known as ‘Machine Learning’ and are part of the broader disciplines of Statistics and Computer Science. Machine Learning techniques enable analysis of large and unstructured datasets and construction of trading strategies. In addition to methods of Classical Machine Learning (that can be thought of as advanced Statistics), there is an increased forcus on investment applications of Deep Learning (an analysis method that relies on multi-layer neural networks), as well as Reinforcement learning (a specific approach that is encouraging algorithms to explore and find the most profitable strategies). While neural networks have been around for decades10, it was only in recent years that they found a broad application across industries. The year 2016 saw the widespread adoption of smart home/mobile products like Amazon Echo11, Google Home and Apple Siri, which relied heavily on Deep Learning algorithms. This success of advanced Machine Learning algorithms in solving complex problems is increasingly enticing investment managers to use the same algorithms.
While there is a lot of hype around Big Data and Machine Learning, researchers estimate that just 0.5% of the data produced is currently being analyzed. These developments provide a compelling reason for market participants to invest in learning about new datasets and Machine Learning toolkits.
As there are quite a lot of terms commonly used to describe Big Data, we provide brief descriptions of Big Data, Machine Learning and Artificial Intelligence below.
Big Data: The systematic collection of large amounts of novel data over the past decade followed by their organization and dissemination has led to the notion of Big Data; see Laney (2001). The moniker ‘Big’ stands in for three prominent characteristics: Volume: The size of data collected and stored through records, transactions, tables, files, etc. is very large; with the subjective lower bound for being called ‘Big’ being revised upward continually. Velocity: The speed with which data is sent or received often marks it as Big Data. Data can be streamed or received in batch mode; it can come in real-time or near-real-time. Variety: Data is often received in a variety of formats – be it structured (e.g. SQL tables or CSV files), semi-structured (e.g. JSON or HTML) or unstructured (e.g. blog post or video message)
Big and alternative datasets include data generated by individuals (social media posts, product reviews, internet search trends, etc.), data generated by business processes (company exhaust data, commercial transaction, credit card data, order book data, etc.) and data generated by sensors (satellite image data, foot and car traffic, ship locations, etc.). The definition of alternative data can also change with time. As a data source becomes widely available, it becomes part of the financial mainstream and is often not deemed ‘alternative’ (e.g. Baltic Dry Index – data from ~600 shipping companies measuring the demand/supply of dry bulk carriers).
Machine Learning (ML): Machine Learning is a part of the broader fields of Computer Science and Statistics. The goal of Machine Learning is to enable computers to learn from their experience in certain tasks. Machine Learning also enables the machine to improve performance as their experience grows. A self-driving car, for example, learns from being initially driven around by a human driver; further, as it drives itself, it reinforces its own learning and gets better with experience. In finance, one can view Machine Learning as an attempt at uncovering relationships between variables, where given historical patterns (input and output), the machine forecasts outcomes out of sample. Machine Learning can also be seen as a model-independent (or statistical or data-driven) way for recognizing patterns in large data sets. Machine Learning techniques include Supervised Learning (methods such as regressions and classifications), Unsupervised Learning (factor analyses and regime identification) as well as fairly new techniques of Deep and Reinforced Learning. Deep learning is based on neural network algorithms, and is used in processing unstructured data (e.g. images, voice, sentiment, etc.) and pattern recognition in structured data.
Artificial Intelligence (AI): Artificial Intelligence is a broader scheme of enabling machines with human-like cognitive intelligence (note that in this report, we sparsely use this term due to its ambiguous interpretation). First attempts of achieving AI involved hardcoding a large number of rules and information into a computer memory. This approach was known as ‘Symbolic AI’, and did not yield good results. Machine Learning is another attempt to achieve AI. Machine Learning and specifically Deep Learning so far represent the most serious attempt at achieving AI. Deep Learning has already yielded some spectacular results in the fields of image and pattern recognition, understanding and translating languages, and automation of complex tasks such as driving a car. While Deep Learning based AI can excel and beat humans in many tasks, it cannot do so in all. For instance, it is still struggling with some basic tasks such as the Winograd Schema Challenge.
