PA1Undergraduate courses
Predictive Analytics I — Machine Learning Tools
The basic paradigm of predictive modeling — classification and prediction — with KNN, Naive Bayes, CART and ensembles applied to real data.
- Introductory
- 4 weeks
- Approx. 15 hours per week
- 3 semester hours
About this course
This course introduces the foundational concepts of predictive analytics and data mining through the two tasks businesses most often need: classification, which assigns a record to a category, and prediction, which estimates a numeric outcome. You partition data so that a model can be judged on records it has not seen, explore and visualize the data before modeling it, then build models with the k-nearest-neighbors, Naive Bayes and classification-and-regression-tree algorithms and measure how well each one performs. The course closes on ensembles — combining several algorithms so that the result is better than any one of them alone.
Who this course is for
Marketing and IT managers, financial analysts and risk managers, accountants, data analysts, data scientists and forecasters. It suits anyone who wants to understand what predictive modeling can do for an organization, run a pilot project without a large setup, or manage a predictive deployment and the technical specialists working on it.
What you will learn
8 outcomes
By the end of this course, you will be able to:
- Visualize and explore data to better understand relationships among variables
- Organize the predictive modeling task and the data flow
- Partition data to provide a basis for assessing predictive models
- Develop machine learning models with the KNN, Naive Bayes and CART algorithms
- Assess the performance of these models with holdout data
- Choose and implement appropriate performance measures for predictive models
- Apply predictive models to generate predictions for new data
- Understand ensemble models and the improvement they bring
Week by week
4-Week curriculum overview
This curriculum is identical across all start dates. Expand a week to explore the topics covered.
- Weeks
- 4
- Your week
- Approx. 15 hours
- Level
- Introductory
PreparationWeek 1
- What is supervised learning
- Data partitioning and holdout samples
- Choosing variables (features)
- Handling missing data
- Visualization and exploration
Classification and predictionWeek 2
- Assessing classification models — confusion matrix, misclassification costs, lift
- Assessing prediction models — common metrics
- K-nearest-neighbors (KNN) — measuring distance
- Choosing k
- Generating classifications and predictions
Bayesian classifiers and CARTWeek 3
- Full Bayes classifier
- Naive Bayes classifier
- Classification and regression trees (CART) — growing the tree
- Avoiding overfit — pruning
- Using trees for classifications and predictions
EnsemblesWeek 4
- Combining multiple algorithms
- Improving results
Instructors
Expert-Led Guidance
Each cohort is led by dedicated instructors and assistant teachers who actively lead weekly discussions, provide personalized feedback and grade your assignments.

Mr. Anthony Babinec
BA in Sociology and MA in Sociology, with a focus on advanced statistics and political sociology, from the University of Chicago. He serves on the editorial board of the Journal of Targeting, Measurement and Analysis for Marketing, and has presented at the AMA's Applied Research Methods Conference, the Advanced Research Techniques Forum, the Sawtooth Software Conference and Statistical Innovation's Statistical Modeling Week.
Before you start
What you need to know first
You will benefit from some familiarity with introductory statistics, specifically regression.
Course Format & Schedule
This is a 4-week, 100% online, asynchronous course.
- No mandatory live sessions: Log in and complete your work at times that fit your schedule.
- Weekly Releases: At the start of each week, you will receive new lecture materials and answer keys for the previous week's exercises.
- Interactive Community: Work through exercises, submit assignments, and engage with your instructor and peers via a private discussion board.
Homework
Short-answer questions testing the concepts and guided data analysis problems using software, alongside supplemental video lectures. There is an end-of-course data modeling project.
Texts
Please choose one of the following textbooks based on the software you plan to use:
- Python: Machine Learning for Business Analytics: Concepts, Techniques, and Applications in Python (2nd ed., 2025) by Shmueli, Bruce, Gedeck, and Patel. Also available at Amazon here.
- R: Machine Learning for Business Analytics: Concepts, Techniques, and Applications in R (2nd ed., 2023) by Shmueli, Bruce, Gedeck, Yahav, and Patel. Also available at Amazon here.
- Analytic Solver Data Mining (previously XLMiner): Machine Learning for Business Analytics: Concepts, Techniques, and Applications in Analytic Solver Data Mining (4th ed., 2023) by Shmueli, Bruce, Deokar, and Patel. Also available at Amazon here.
The same text is also used in Predictive Analytics II — Neural Nets and Regression and Predictive Analytics III — Dimension Reduction, Clustering and Association Rules. So one copy covers all three courses.
Software
This is a hands-on course in which you will apply data mining algorithms to real datasets. The course can be completed using Python or R, both of which are free, open-source programming languages. Corresponding editions of the course text are available for Python and R, making these the recommended options for completing the course without additional software costs. Worked examples are also available in Analytic Solver Data Mining (ASDM), an add-in for Microsoft Excel. If you choose this option, you will need both Microsoft Excel and ASDM. Course participants will receive a license for ASDM for nominal cost — this is a special version for this course. IMPORTANT: Do NOT download the free trial version available at solver.com
FAQ
Do I need to be advanced in Excel?This course
No. You only need basic Excel skills. The algorithms are executed using an Excel add-in rather than complex formulas. The course focuses on selecting variables, evaluating models, and interpreting outputs—not advanced spreadsheet skills.
Is this a machine learning course or a statistics course?This course
It is machine learning designed for business analytics and real-world decision-making. You will cover core machine learning algorithms—such as k-nearest neighbors, Naive Bayes, decision trees, and ensembles—with a heavy emphasis on data partitioning, holdout validation, and performance measurement. We focus on building models that generalize to new data, not just training data.
Is there a money-back or satisfaction guarantee?Every course
We handle cancellation and refund requests on an individual basis. If a course is not meeting your expectations or your circumstances change, please reach out to our team (support@learnstatistics.org) so we can work with you on a solution.
Can I transfer to a later start date or withdraw after the course begins?Every course
Yes. If you need to defer your enrollment to a future start date, withdraw mid-course, or transfer your seat to a colleague, please contact us directly. We address these requests flexibly on a case-by-case basis.
Who teaches the course?Every course
Each cohort is led by dedicated instructors and assistant teachers who actively lead weekly discussions, provide personalized feedback and grade your assignments.
Something not answered here? Ask us before you book — and the terms set out what a purchase covers.
Credit, and where it counts
This course carries a verified figure of 3 semester hours. A credit recommendation from the American Council on Education says what a course is worth; the institution receiving it decides whether to award it.
Anyone can take this subject: there is no application, and you do not have to be studying at a university. It also counts toward a degree at Thomas Edison State University, which decides what to award for it. The full position — including the limits each university publishes is on its own page, and worth reading before you buy.