Feature Selection and Noise Control in High-Dimensional Data: Tourism vs. Healthcare
DOI:
https://doi.org/10.63313/AJET.9063Keywords:
High-dimensional data, feature selection, noise control, multi-source governance, decision efficiency, smart tourism, precision healthcare, data purification, decision usabilityAbstract
This study proposes a unified governance-oriented framework for managing high-dimensional multi-source data through the integration of feature selection and noise control, with the explicit goal of improving decision efficiency rather than purely technical accuracy. Drawing on comparative evidence from smart tourism and healthcare systems, the research demonstrates that the main challenge of high-dimensional data lies not in domain-specific content but in shared governance structures, including multi-source integration, systematic noise generation, and decision pressure under uncertainty. Feature selection is conceptualized as an institutional information purification mechanism that filters variables according to relevance, interpretability, and actionability. Noise control is framed as a governance pipeline involving profiling, harmonization, cleaning, validation, and continuous monitoring. Together, these processes transform complex and heterogeneous datasets into reliable and governable decision inputs. The study shifts evaluation criteria from medical or predictive accuracy toward governance-oriented indicators such as latency, stability, interpretability, robustness, and actionability. Tourism is positioned as the methodological origin and practical field of governance-oriented data purification, while healthcare serves as a comparative application context. The findings establish feature selection and noise control as a transferable management science methodology for high-dimensional data governance and institutional decision-making.
References
[1] Adamış, E., & Pınarbaşı, F. (2022). Unfolding visual characteristics of social media communication: reflections of smart tourism destinations. Journal of Hospitality and Tourism Technology, 13(1), 34-61.
[2] Batini, C., & Scannapieco, M. (2016). Data and information quality. Cham, Switzerland: Springer International Publishing, 63.
[3] Bellman, R. (1961, January). A mathematical formulation of variational processes of adaptive type. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics (Vol. 4, pp. 37-49). University of California Press.
[4] Borgman, C. L. (2017). Big data, little data, no data: Scholarship in the networked world. MIT press.
[5] Bowker, G. C., & Star, S. L. (2000). Sorting things out: Classification and its consequences. MIT press.
[6] Breiman, L. (1996). Heuristics of instability and stabilization in model selection. The annals of statistics, 24(6), 2350-2383.
[7] Chen, M., Mao, S., & Liu, Y. (2014). Big data: A survey. Mobile networks and applications, 19(2), 171-209.
[8] Davenport, T. H., & Harris, J. G. (2007). Competing on analytics: the new science of Winning. Harvard business review press, Language, 15(217), 24.
[9] Donabedian, A. (1988). The quality of care: how can it be assessed?. Jama, 260(12), 1743-1748.
[10] Fellegi, I. P., & Sunter, A. B. (1969). A theory for record linkage. Journal of the American statistical association, 64(328), 1183-1210.
[11] Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., ... & Vayena, E. (2018). AI4People-An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and machines, 28(4), 689-707.
[12] Friedman, J. H., Hastie, T., & Tibshirani, R. (2010). Regularization paths for generalized linear models via coordinate descent. Journal of statistical software, 33, 1-22.
[13] Fuller, W. A. (2009). Measurement error models. John Wiley & Sons.
[14] Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of machine learning research, 3(Mar), 1157-1182.
[15] Hall, M. A. (1999). Correlation-based feature selection for machine learning (Doctoral dissertation, The University of Waikato).
[16] Hamilton, J. D. (1989). A new approach to the economic analysis of nonstationary time series and the business cycle. Econometrica: Journal of the econometric society, 357-384.
[17] Heckman, J. (2013). Sample selection bias as a specification error. Applied Econometrics, 31(3), 129-137.
[18] Herland, M., Khoshgoftaar, T. M., & Wald, R. (2014). A review of data mining using big data in health informatics. Journal of Big data, 1(1), 2.
[19] Huber, P. J., & Ronchetti, E. M. (2009). Robust statistics. 2nd john wiley & sons. Hoboken, NJ, 2.
[20] Janssen, M., & Van Der Voort, H. (2016). Adaptive governance: Towards a stable, accountable and responsive government. Government Information Quarterly, 33(1), 1-5.
[21] Kahn, M. G., Brown, J. S., Chun, A. T., Davidson, B. N., Meeker, D., Ryan, P. B., ... & Zozus, M. N. (2015). Transparent reporting of data quality in distributed data networks. Egems, 3(1), 1052.
[22] Kohavi, R., & John, G. H. (1997). Wrappers for feature subset selection. Artificial intelligence, 97(1-2), 273-324.
[23] Leek, J. T., Scharpf, R. B., Bravo, H. C., Simcha, D., Langmead, B., Johnson, W. E., ... & Irizarry, R. A. (2010). Tackling the widespread and critical impact of batch effects in high-throughput data. Nature Reviews Genetics, 11(10), 733-739.
[24] Lipton, Z. C. (2018). The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16(3), 31-57.
[25] Liu, H., & Motoda, H. (2012). Feature selection for knowledge discovery and data mining (Vol. 454). Springer science & business media.
[26] Luca, M., & Zervas, G. (2016). Fake it till you make it: Reputation, competition, and Yelp review fraud. Management science, 62(12), 3412-3427.
[27] March, J. G., & Simon, H. A. (1993). Organizations. John wiley & sons.
[28] Mintzberg, H., Raisinghani, D., & Theoret, A. (1976). The structure of" unstructured" decision processes. Administrative science quarterly, 246-275.
[29] Pipino, L. L., Lee, Y. W., & Wang, R. Y. (2002). Data quality assessment. Communications of the ACM, 45(4), 211-218.
[30] Power, D. J., & Heavin, C. (2017). Decision support, analytics, and business intelligence. Business Expert Press.
[31] Raghupathi, W., & Raghupathi, V. (2014). Big data analytics in healthcare: promise and potential. Health information science and systems, 2(1), 3.
[32] Rahm, E., & Do Hong, H. (2000). Data cleaning: Problems and current approaches.
[33] Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581-592.
[34] Shmueli, G., & Koppius, O. R. (2011). Predictive analytics in information systems research. MIS quarterly, 553-572.
[35] Simon, H. A. (1960). The new science of management decision.
[36] Suchman, M. C. (1995). Managing legitimacy: Strategic and institutional approaches. Academy of management review, 20(3), 571-610.
[37] Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of management information systems, 12(4), 5-33.
[38] Zhu, X., & Wu, X. (2004). Class noise vs. attribute noise: A quantitative study. Artificial intelligence review, 22(3), 177-210.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by author(s) and Erytis Publishing Limited

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.













