Detecting periods of abnormal market behaviour in the S&P 500 with GARCH, ARIMA, and Isolation Forest, plus clustering, feature selection, and next-day direction classifiers.
Solo project · CMPT 459 Data Mining (Prof. Martin Ester), Simon Fraser University · Fall 2024 With the instructor's approval, the project also brings in time-series methods from STAT 485 (ARIMA/GARCH), which I learned in R and implemented here in Python.
Stack: Python · pandas · NumPy · statsmodels · arch · scikit-learn · XGBoost · matplotlib · seaborn
- Found and worked around a data-quality problem. In the newest release of the source dataset, many constituents had no price history at all, including heavily weighted names such as AAPL, AMZN, and NVDA. About 333 tickers were missing, roughly 69% of index weight. After ruling out my own code as the cause, I rolled back to an earlier complete release (version 960).
- Compared three anomaly detectors on a time-based split: fit on data before 2019, then evaluate on 2019 onward. Isolation Forest and ARIMA both flag the March 2020 COVID crash on held-out data. GARCH catches it only when fit on the full series, which says more about how the forecast was set up than about the model itself.
- Tested whether simple features predict the index's next-day direction. They don't. Every classifier landed close to the majority-class baseline, which is the expected result for daily market data.
| Stage | What I did |
|---|---|
| Preprocessing | Daily returns from closing prices; standardized returns for GARCH and Isolation Forest; differencing for ARIMA stationarity; time-based train/test split at 2019-01-01 |
| EDA | Correlation heatmaps (top 15 weighted stocks and sector averages vs. the index), sector weights, and a missing-data audit across dataset versions |
| Anomaly detection | GARCH: flag days where conditional volatility exceeds mean + 2σ. ARIMA(1,1,1): flag residuals outside a rolling standard-deviation band. Isolation Forest: 5% contamination on daily returns |
| Volatility forecasting | 7-day GARCH(2,2) vs. ARCH(2) forecasts |
| Clustering | K-Means (with DBSCAN explored) on per-stock return and volume |
| Feature selection | RFE, L1-regularized logistic regression, mutual information |
| Classification | k-NN, logistic regression, Random Forest, XGBoost, SVM on aggregated daily features (volume, mean return, cross-sectional volatility, 5-day moving average, 5-day momentum) |
| GARCH conditional volatility (full series) | Isolation Forest on held-out returns (2019+) |
|---|---|
![]() |
![]() |
- Isolation Forest gave the clearest held-out result. Its flags cluster tightly around March 2020 and through the 2022 drawdown.
- GARCH picks out the 2020 and 2022 volatility spikes when fit on the full series. When fit only on pre-2019 data and projected forward, its forecast settles to a long-run level within weeks, so on the test period it flags only early January 2019 (the red points in the consolidated plot). Re-estimating the model on a rolling window would fix this.
- ARIMA residuals catch the major events but also produce scattered false positives, so threshold selection matters.
The test set has 504 days, of which 52.4% were up days. A model that always predicts "up" therefore scores 52.4%.
| Model | Test accuracy |
|---|---|
| k-NN (k=5) | 47.8% |
| Random Forest | 50.6% (ROC-AUC 0.50) |
| XGBoost | 49.2% |
| SVM (RBF) | 52.0% |
| Logistic regression | 54.0% |
| Logistic regression, L1-selected features | 53.4% |
None of the models meaningfully beat the baseline. The logistic and SVM models reach their scores mostly by predicting "up" nearly every day, with recall of 0.95 or higher on the up class. My read is that these features carry little next-day signal. That is consistent with how noisy daily index returns are, not evidence that a better model is waiting to be found.
- Audit the data before building models. I built most of the model pipeline against the index file alone and found the missing-constituent problem late, which cost time.
- Target leakage in an early experiment.
main1.1.ipynbincludes a Random Forest that scores 100% accuracy. That result is leakage: the label (Return > 0) was derived from a feature that was also in the input. It is superseded by the leakage-free setup inPosterCode.ipynb, where the label is the next day's index direction. - The train/test split is random, not chronological, for the classifiers. A walk-forward split would be the more honest evaluation for time-series data.
- Future work: rolling GARCH re-estimation, walk-forward validation, and macro or sector-level features for the anomaly models.
├── main1.1.ipynb # Main notebook: preprocessing, EDA, anomaly detection
├── PosterCode.ipynb # Clustering, feature selection, classifiers (results above)
├── ID_test.ipynb # Dataset investigation
├── Investigate Datasets.py # EDA plots: top-15 stocks, sector weights, heatmaps
├── Project Report.pdf # Full written report
├── Poster.pptx # Project poster
├── Project Proposal/ # Original proposal
├── Project plots and versions/ # Generated figures
├── Other code/ # Earlier iterations
├── sp500_index.csv, sp500_companies.csv
└── Working Data Set - V960.zip # Contains sp500_stocks.csv (too large to commit unzipped)
pip install numpy pandas matplotlib seaborn statsmodels arch scikit-learn xgboost scipy tslearn yfinance
unzip "Working Data Set - V960.zip" # move sp500_stocks.csv next to the notebooksThen run main1.1.ipynb for the anomaly detection and PosterCode.ipynb for the classification experiments.
The full write-up is in Project Report.pdf.


