Reinforcement Learning

Atari

An especially promising approach to Machine Learning is reinforcement learning. The goal of reinforcement learning is to choose a course of actions that maximizes some reward. For instance, one may look for a set of trading rules that maximizes PnL after 100 trades.

Unlike supervised learning (which is typically a one-step process), the model doesn’t know the correct action at each step, but learns over time which succession of steps led to the highest reward at the end of the process.

A fascinating illustration is a performance of the reinforcement learning algorithm in playing a simple Atari videogame (Google’s DeepMind). After training the algorithm on this simple task, the machine can easily beat a human player. While most human players (and similarly traders) learn by maximizing rewards, humans tend to stop refining a strategy after a certain level of performance is reached. On the other hand, the machine keeps on refining, learning, and improving performance until it achieves perfection.

At the core of reinforcement learning are two challenges that the algorithm needs to solve:

  1. Explore vs. Exploit dilemma – should the algorithm explore new alternative actions that may not be immediately optimal but may maximize the final reward (or stick to the established ones that maximize the immediate reward);
  2. Credit assignment problem – given that we know the final reward only at the last step (e.g. end of game, final PnL), it is not straightforward to assess which step during the process was critical for the final success. Much of the reinforcement learning literature aims to answer the twin questions of the credit assignment problem and exploration-exploitation dilemma.

When used in combination with Deep Learning, reinforcement learning has yielded some of the most prominent successes in machine learning, such as self-driving cars. Within finance, reinforcement learning already found application in execution algorithms and higher-frequency systematic trading strategies.
Reinforcement learning has attributes of both supervised and unsupervised learning. In supervised learning, we have access to a training set, where the correct output “y” is known for each input. At the other end of the spectrum was unsupervised learning where we had no correct output “y” and we are learning the structure of data. In reinforcement learning we are given a series of inputs and we are expected to predict y at each step. However, instead of getting an instantaneous feedback at each step, we need to study different paths/sequences to understand which one gives the optimal final result.

Ads Blocker Image Powered by Code Help Pro

Ads Blocker Detected!!!

We have detected that you are using extensions to block ads. Please support us by disabling these ads blocker.

Powered By
100% Free SEO Tools - Tool Kits PRO