Johnny Huincahue
Bachelor’s Thesis
Thesis Advisor
- Héctor Olivero
Profesor Co-Guía
- Karine Bertin
- Cristian Meza
Resumen
This study evaluates an approach for detecting stock market manipulation by combining supervised machine learning and—where temporal resolution permits—Functional Data Analysis. Two scenarios are examined. For real daily data (the FLC Group case), observations are represented by tabular variables derived from returns, volume, and volatility; classical classifiers are compared under conditions of severe class imbalance using metrics focused on the minority class (primarily the F2-score and the area under the Precision-Recall curve). In parallel, synthetic intraday data are generated on a fixed 5-second grid using three models (MBG, Heston, and Cont-Müller), incorporating "pump-and-dump" manipulation episodes of varying intensity and imbalance levels. This enables the application of functional smoothing (B-splines/wavelets), Functional Principal Component Analysis, and classification techniques, including the functional k-Nearest Neighbors algorithm. The results demonstrate that explicitly addressing class imbalance is crucial to preventing the model from collapsing toward the majority class. Furthermore, detectability improves when manipulation is more intense or when variables providing a direct signal of the mechanism (such as market depth in the Cont-Müller model) are available. A key limitation is that the functional analysis of real data is constrained by the lack of intraday information; consequently, future work will aim to incorporate real high-frequency data and calibrate simulations to better approximate realistic market conditions.