Anomaly detection in commitment of traders (COT) report data using manifold learning approach

Document Type : Original Article

Authors

1 Department of Mathematics, Faculty of Mathematics and Computer Science, Iran University of Science and Technology, Tehran, Iran.

2 Department of Computer Science, Faculty of Mathematics and Computer Science, Iran University of Science and Technology, Tehran, Iran.

Abstract
Purpose: The objective of this research is to present a comprehensive framework for identifying anomalies in Commitments of Traders (COT) report data and to investigate their role in detecting economic trends, market disruptions, and sudden changes.
Methodology: In this study, following the preprocessing of the COT data, statistical methods, including the standard score and the interquartile range, were combined with machine learning algorithms, including Isolation Forest and One-Class Support Vector Machine. Then, by creating ensemble anomalies, the points identified as anomalous by all methods were extracted. Subsequently, linear and non-linear dimensionality reduction methods PCA, Isomap, UMAP, and LLE, were applied to the COT dataset, and the One-Class Support Vector Machine, Local Outlier Factor, Isolation Forest, and K-Means algorithms were implemented for anomaly detection.
Findings: The findings demonstrated that the detected anomalies closely coincide with prominent macroeconomic disruptions, notably the 2008 Global Financial Crisis and the economic shock of the 2020 COVID-19 pandemic. Additionally, the consensus framework, grounded in the intersection of model outputs, effectively filtered alarms driven by algorithm-specific sensitivities, thereby yielding a focused subset of 1,638 observations characterized by the highest inter-algorithm consensus.
Originality/Value: This research, by integrating statistical methods, machine learning, and Manifold Learning, provides an accurate and reliable framework for analyzing complex financial market data and creates the foundation for developing interactive dashboards for real-time market monitoring.

Keywords


[1]     Irwin, S. H., & Sanders, D. R. (2011). Index funds, financialization, and commodity futures markets. Applied economic perspectives and policy, 33(1), 1–31. https://doi.org/10.1093/aepp/ppq032
[2]     Brooks, C., & Wichmann, R. (2019). Eviews guide to accompany introductory econometrics for finance. SSRN. https://ssrn.com/abstract=3466908
[3]     Liu, F. T., Ting, K. M., & Zhou, Z. H. (2008). Isolation forest. In 2008 eighth ieee international conference on data mining (pp. 413-422). IEEE. https://doi.org/10.1109/ICDM.2008.17
[4]     Sato, Y. (2025). Shortage, liquidity, and the effectiveness of monetary policy: Evidence from the us commodity futures markets. Available at SSRN 5567658. https://dx.doi.org/10.2139/ssrn.5567658
[5]     Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM computing surveys (CSUR)41(3), 1-58. https://doi.org/10.1145/1541880.1541882
[6]     Lensen, A., Zhang, M., & Xue, B. (2020). Multi-objective genetic programming for manifold learning: Balancing quality and dimensionality. Genetic programming and evolvable machines, 21(3), 399–431. https://doi.org/10.1007/s10710-020-09375-4
[7]     Kritzman, M., & Li, Y. (2010). Skulls, financial turbulence, and risk management. Financial analysts journal, 66(5), 30–41. https://doi.org/10.2469/faj.v66.n5.3
[8]     Phatak, A. A., Wieland, F. G., Vempala, K., Volkmar, F., & Memmert, D. (2021). Artificial intelligence based body sensor network framework—narrative review: proposing an end-to-end framework using wearable sensors, real-time location systems and artificial intelligence/machine learning algorithms for data collection, data mining and knowledge discovery in sports and healthcare. Sports Medicine-Open7(1), 79. https://doi.org/10.1186/s40798-021-00372-0
[9]     Liu, F. T., Ting, K. M., & Zhou, Z. H. (2012). Isolation-based anomaly detection. ACM transactions on knowledge discovery from data (TKDD)6(1), 1-39. https://doi.org/10.1145/2133360.2133363
[10]   Tax, D. M., & Duin, R. P. (2004). Support vector data description. Machine learning54(1), 45-66. https://doi.org/10.1023/B:MACH.0000008084.60811.49
[11]   Aulerich, N. M., Irwin, S. H., & Garcia, P. (2014). Bubbles, food prices, and speculation: Evidence from the CFTC's daily large trader data files. In The economics of food price volatility (pp. 211-253). University of Chicago Press. https://www.nber.org/system/files/chapters/c12814/c12814.pdf