{"id":2068,"date":"2019-06-28T11:57:44","date_gmt":"2019-06-28T06:27:44","guid":{"rendered":"http:\/\/blog.tenthplanet.in\/?p=2068"},"modified":"2026-07-03T14:45:40","modified_gmt":"2026-07-03T09:15:40","slug":"extracting-insights-data-exploratory-data-analysis","status":"publish","type":"post","link":"https:\/\/tenthplanet.in\/blogs\/extracting-insights-data-exploratory-data-analysis\/","title":{"rendered":"Extracting Meaningful Insights from Data using Exploratory Data Analysis"},"content":{"rendered":"<h3>Introduction<\/h3>\n<p>Exploratory Data Analysis (EDA) is a set of techniques used for data exploration. It identifies interesting characteristics, relationships and patterns, key attributes and anomalies. EDA is facilitated through graphical and statistical analysis.<\/p>\n<h3>Demonstration<\/h3>\n<p>I would like to share my understanding through an example EDA that I have performed on <a href=\"https:\/\/archive.ics.uci.edu\/ml\/datasets\/wine+quality\" target=\"_blank\" rel=\"noopener\">Wine quality data set<\/a>, obtained from the UCI Machine Learning repository. The following are few of the insights I have derived from the analysis. The <strong>objective<\/strong> of the analysis is to understand how each factor influences the rating of wine.<\/p>\n<p>EDA involves the following steps:<\/p>\n<ul>\n<li>Calculating mean, median, mode and quartile values of the data in each column of the data set.<\/li>\n<li>Visualising the data and observing how the values are distributed.<\/li>\n<li>Identifying outliers (data that do not fall within the ambit of our analysis).<\/li>\n<li>The range of data in each column is calculated, which can show us the magnitude of values present as well as the minimum and maximum values.<\/li>\n<\/ul>\n<p><strong>Data Preparation &amp; EDA<\/strong><\/p>\n<p>Data is collected from various sources and loaded to data warehouse using Pentaho+ Data Integration. The Pentaho Plus Data Integration Tool performs the cleansing, transformation, applying rules and stores in data warehouse.<\/p>\n<p>This data is further used for the exploratory data analysis, creating data pipeline and building model.<\/p>\n<p><strong>Uni-variate plots<\/strong> help in understanding the nature of values in each of the attributes.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-2071\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog2-1-300x214.png\" alt=\"\" width=\"330\" height=\"235\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog2-1-300x214.png 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog2-1-1024x731.png 1024w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog2-1-768x549.png 768w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog2-1.png 1344w\" sizes=\"auto, (max-width: 330px) 100vw, 330px\" \/><\/p>\n<p><strong>Line plots for the variables in the data<\/strong><\/p>\n<p>From the above line plots, we can infer that pH and density have normal distribution, whereas SO2, sulphates and alcohol have skewed distribution. Quality shows comb distribution, which means that certain data points may have been rounded off to the nearest whole number.<\/p>\n<p><strong>Bi-variate plots<\/strong> help in understanding the relationship between two plotted attributes.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-2072\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog4-1-300x214.png\" alt=\"\" width=\"338\" height=\"241\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog4-1-300x214.png 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog4-1-1024x731.png 1024w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog4-1-768x549.png 768w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog4-1.png 1344w\" sizes=\"auto, (max-width: 338px) 100vw, 338px\" \/><\/p>\n<p><strong>\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0Quality Vs Alcohol scatter plot<\/strong><\/p>\n<p>From the above scatter plot, we can infer that good-quality wines usually have higher alcohol content than low-quality wines .<\/p>\n<p><strong>Correlation Plot<\/strong> shows linear correlation between two variables.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-2073\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog5-1-300x214.png\" alt=\"\" width=\"393\" height=\"280\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog5-1-300x214.png 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog5-1-1024x731.png 1024w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog5-1-768x549.png 768w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog5-1.png 1344w\" sizes=\"auto, (max-width: 393px) 100vw, 393px\" \/><\/p>\n<p><strong>\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0Correlation plot<\/strong><\/p>\n<p>From the above correlation plot, we can list the attributes that are strongly correlated:<\/p>\n<p>Quality and Alcohol (positively correlated)<br \/>\nDensity and Acidity (positively correlated)<br \/>\nDensity and Alcohol (negatively correlated)<br \/>\nFree SO2 and Total SO2 (positively correlated)<br \/>\nAcidity and pH (negatively correlated)<\/p>\n<p><strong>Multivariate Plots<\/strong> find relationships between more than 2 variables. They become difficult to interpret when the number of variables is greater than 3.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-2074\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog6-1-300x214.png\" alt=\"\" width=\"390\" height=\"278\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog6-1-300x214.png 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog6-1-1024x731.png 1024w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog6-1-768x549.png 768w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2019\/06\/blog6-1.png 1344w\" sizes=\"auto, (max-width: 390px) 100vw, 390px\" \/><\/p>\n<p><strong>\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a03-D scatter plot<\/strong><\/p>\n<p>From the above graph, we can infer that most good-quality wines contain less amount of chlorides and an acceptable density.<\/p>\n<h3><strong>Conclusion<\/strong><\/h3>\n<p>Various types of plots and techniques are available to represent and analyse your data like <a href=\"https:\/\/en.wikipedia.org\/wiki\/Box_plot\" target=\"_blank\" rel=\"noopener\">Box plot<\/a>, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Histogram\" target=\"_blank\" rel=\"noopener\">Histogram<\/a>, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Run_chart\" target=\"_blank\" rel=\"noopener\">Run chart<\/a>, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Pareto_chart\" target=\"_blank\" rel=\"noopener\">Pareto char<\/a>t, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Scatter_plot\" target=\"_blank\" rel=\"noopener\">Scatter plot<\/a>, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Odds_ratio\" target=\"_blank\" rel=\"noopener\">Odds ratio<\/a> and more. It is crucial to utilise the right plot so as to grab the essence of pattern from the data.<\/p>\n<p>Therefore EDA is the process of obtaining any possible intuition regarding the data in the beginning itself. Skipping this process might lead to skewed data with too many outliers and missing values. This, in turn, could generate inaccurate models, lead to choosing wrong variables for the model or the wrong model itself.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Raw data is normally ambiguous and difficult to interpret. Cleaning it is essential in order to understand the relationships between the variables present in the data.<\/p>\n","protected":false},"author":23,"featured_media":2080,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[424],"tags":[484,468,489,478],"class_list":["post-2068","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-pentaho","tag-data-cleaning","tag-data-science","tag-data-visualization","tag-exploratory-data-analysis"],"acf":[],"_links":{"self":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts\/2068","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/users\/23"}],"replies":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/comments?post=2068"}],"version-history":[{"count":1,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts\/2068\/revisions"}],"predecessor-version":[{"id":11433,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts\/2068\/revisions\/11433"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/media\/2080"}],"wp:attachment":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/media?parent=2068"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/categories?post=2068"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/tags?post=2068"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}