Author(s):
Renu, Sangeeta, Y.K. Gupta
Email(s):
ykgbkbiet123@gmail.com , yk.gupta@bkbiet.ac.in , ykgbkbiet@rediffmail.com
DOI:
10.52711/2321-581X.2026.00002
Address:
Renu1, Sangeeta1, Y.K. Gupta2*
1Research Scholar, B K Birla Institute of Engineering and Technology, Pilani, Rajasthan, India.
2Professor and Head Department of Chemistry, B K Birla Institute of Engineering and Technology, Pilani, Rajasthan, India.
*Corresponding Author
Published In:
Volume - 17,
Issue - 1,
Year - 2026
ABSTRACT:
This paper offers a thorough analysis of the classification of iris flower species using machine learning methods. Through data investigation and visualization, the popular Iris dataset, which comprises 150 samples and four essential features like sepal length, sepal width, petal length, and petal width was examined. Pairplots, boxplots, and correlation heatmaps were used in exploratory data analysis (EDA) to show distinct patterns and separability across the three species. Some learning methods, including Logistic Regression, K-Nearest Neighbors, Support Vector Machine, and Decision Tree were used to create classification models once the data had been pre-processed. Accuracy scores and confusion matrices were used to assess these models performance. The Support Vector Machine (SVM) produced the most accurate and dependable classification outcomes out of all the models. This work demonstrates that conventional machine learning algorithms can effectively categorize Iris species with appropriate analysis and model selection, making this dataset a perfect benchmark for novices and machine learning researchers.
Cite this article:
Renu, Sangeeta, Y.K. Gupta. An Exploratory Data Analysis and Classification Study on the Iris Dataset. Research Journal of Engineering and Technology. 2026;17(1):16-4. doi: 10.52711/2321-581X.2026.00002
Cite(Electronic):
Renu, Sangeeta, Y.K. Gupta. An Exploratory Data Analysis and Classification Study on the Iris Dataset. Research Journal of Engineering and Technology. 2026;17(1):16-4. doi: 10.52711/2321-581X.2026.00002 Available on: https://rjetonline.com/AbstractView.aspx?PID=2026-17-1-2
7. REFERENCES:
1. Fisher, R. A. The use of multiple measurements in taxonomic problems. Annals of Eugenics. 1936; 7(2): 179–188.
2. Bishop, C. M. Pattern Recognition and Machine Learning. Springer. 2006; 1(2): 79–88.
3. Cortes, C., and Vapnik, V. Support-vector networks. Machine Learning, 1995; 20(3): 273–297.
4. Pedregosa, F., Varoquaux, G., Gramfort, A., et al. 2011; 13(3): 373–397.
5. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2009; 20(3): 273–297
6. Breiman, L. Random forests. Machine Learning, 2001; 45(1): 5–32.
7. Quinlan, J. R. Induction of decision trees. Machine Learning. 1986; 1: 81–106.