Share with your friends










Submit

Analytics Magazine

Data Mining: Predicting stock price movements

January/February 2011

CLICK HERE TO GO TO THE DIGITAL VERSION OF THIS ARTICLE

INFORMS data mining contest attracts 894 participants representing 147 teams from 27 countries.

By Durai Sundaramoorthi, Philipp e Belanger and Louis Duclos-Gosselin

Cole Harris of Exagen Diagnostics (www.exagen.com) won the 2010 Data Mining Contest that required participants to develop a predictive analysis solution to predict stock price movement (increase or decrease) in “the next 60 minutes” in five-minute intervals. (For example, at 9:30 a.m. will XYZ’s stock price increase or decrease at 10:30 a.m. based on historical data?)

Christopher Hefele of AT&T finished second and Nan Zhou from the University of Pittsburgh placed third in the contest organized and sponsored by the INFORMS Data Mining Section. The contest drew 894 participants from 147 teams representing 27 countries, making it the largest event of its kind in the world.

Day traders, mutual fund traders and hedge funds have always tried to predict the direction of stock prices in the next few hours. Predictive analysis solutions developed in the contest aim to pursue this objective. Hedge funds can use these types of solutions to build complex strategies to be executed automatically. Mutual fund traders can use them to achieve a “best execution” of fund managers’ buy/sell orders, while day traders can use them to realize fast profits over short periods of time.

The predictive analysis solutions produce valuable business results in the real world applications by building recommendations systems.

Here’s how it works:

First, the data miners develop the data mining techniques by:

  • making the different databases communicate with each other to collect all the information on the back office databases;
  • extracting the needed data and dealing with missing data and outliers;
  • preparing the data for the data mining algorithm;
  • writing the code relative to the data mining algorithm in the company systems;
  • predicting the data from the test database, to validate the performance of the data mining techniques; and
  • writing the code relative to the implementation of all the processes on the company systems.

To verify that the data mining algorithm will work when they implement it, the “miners” predict (score) the test database observations and validate the performance. The data mining algorithm should have similar performance in the training database and in the test database to be considered a “superior performer.” The data mining algorithm with the best performance is implemented. The test database generally represents 30 percent of the database.

When the implementation of the data mining techniques is complete, one code runs each five minutes. The code collects, from the database, the different information needed, applies the data mining algorithm, predicts the studied stock price will increase or decrease in 60 minutes and recommends the appropriate action to maximize the profit. Figure 1 illustrates this process.

In the INFORMS Data Mining Contest, participants were provided with a set of macro-economic and high frequency financial data to build their predictive analysis solutions. The data were composed of stock prices, sector indexes, economic indicators and expert predictions on economic indicators. The database was separated into two data sets: the training set for building predictive analysis model(s) and the test set (in which the target variable has been excluded) for evaluating participants’ predictions. The participants built their predictive analysis solutions using the training set and implemented it on the test set by predicting target variable.

The top three finishers in the contest presented their methods at the 2010 INFORMS Annual Meeting in Austin, Texas, in November. To view slides of their solutions, see: www.kaggle.com/informs2010.


Durai Sundaramoorthi is an assistant professor at Missouri Western State University, Philippe Belanger of Laval University served as co-chair of the 2010 Data Mining Contest, and Louis Duclos-Gosselin (louis.gosselin@hotmail.com) of Sinapse chaired the contest, a role he will continue in 2011.

CLICK HERE TO GO TO THE DIGITAL VERSION OF THIS ARTICLE



Headlines

Meet CIMON, the first AI-powered astronaut assistant

CIMON, the world’s first artificial intelligence-enabled astronaut assistant, made its debut aboard the International Space Station. The ISS’s newest crew member, developed and built in Germany, was called into action on Nov. 15 with the command, “Wake up, CIMON!,” by German ESA astronaut Alexander Gerst, who has been living and working on the ISS since June 8. Read more →

Yale research on immigration, aging runners makes news

A recent study by Yale University professor and former INFORMS President Edward H. Kaplan (photo) and Yale colleague Jonathan Feinstein and Mohammad M. Fazel-Zarandi of MIT suggests that the number of undocumented immigrants in the United States is nearly twice as many as experts previously thought. Since its publication last month, the study, which estimates the number of such immigrants at 22.1 million instead of 11.3 million, has garnered worldwide attention from major media outlets including the Los Angeles Times, the Boston Globe, Fox News, Bloomberg News and the Daily Mail. Read more →

New salary survey paints optimistic picture for analytics professionals

Harnham, a global leader in data and analytics recruitment, recently released the 2018 editions of its salary guides for the United Kingdom, the United States and Europe. Having heard from thousands of data and analytics professionals across the globe, Harnham has gained an invaluable insight into key industry salaries and trends across a wide variety of analytics specialties and sectors. Read more →

UPCOMING ANALYTICS EVENTS

INFORMS-SPONSORED EVENTS

Winter Simulation Conference
Dec. 9-12, 2018, Gothenburg, Sweden

INFORMS Computing Society Conference
Jan. 6-8, 2019; Knoxville, Tenn.

INFORMS Conference on Business Analytics & Operations Research
April 14-16, 2019; Austin, Texas

INFORMS International Conference
June 9-12, 2019; Cancun, Mexico

INFORMS Marketing Science Conference
June 20-22; Rome, Italy

INFORMS Applied Probability Conference
July 2-4, 2019; Brisbane, Australia

INFORMS Healthcare Conference
July 27-29, 2019; Boston, Mass.

2019 INFORMS Annual Meeting
Oct. 20-23, 2019; Seattle, Wash.

Winter Simulation Conference
Dec. 8-11, 2019: National Harbor, Md.

OTHER EVENTS

Applied AI & Machine Learning | Comprehensive
Dec. 3, 2018 (live online)


Advancing the Analytics-Driven Organization
Jan. 28–31, 2019, 1 p.m.– 5 p.m. (live online)

CAP® EXAM SCHEDULE

CAP® Exam computer-based testing sites are available in 700 locations worldwide. Take the exam close to home and on your schedule:


 
For more information, go to 
https://www.certifiedanalytics.org.