{"id":244,"date":"2018-02-14T10:39:02","date_gmt":"2018-02-14T05:09:02","guid":{"rendered":"http:\/\/blog.tenthplanet.in\/?p=244"},"modified":"2026-03-03T10:15:38","modified_gmt":"2026-03-03T10:15:38","slug":"matching-words-which-are-similar-phonetically","status":"publish","type":"post","link":"https:\/\/tenthplanet.in\/blogs\/matching-words-which-are-similar-phonetically\/","title":{"rendered":"Matching Words Which are Similar Phonetically &#8211; Fuzzy Logic"},"content":{"rendered":"<h3 class=\"western\">Introduction<\/h3>\n<h4 class=\"western\">Fuzzy Logic:<\/h4>\n<p>Fuzzy matching is a method that provides an improved ability to process word-based matching queries to find matching phrases or sentences from a database.\u00a0When an exact match is not found for a sentence or phrase, fuzzy matching can be applied.<\/p>\n<p>Fuzzy matching attempts to find a match which, although not a 100 percent match, is above the threshold matching percentage set by the application.<\/p>\n<p>It works with matches that may be less than 100% perfect when finding correspondences between segments of a text and entries in a database of previous translations.<\/p>\n<h3 class=\"western\">Business Situation<\/h3>\n<p>Customer having the practitioner&#8217;s prescription for the patient&#8217;s diagnosis and diseases in a text file, collected from the various hospitals and data has entered manually. Now the problem is to map the practitioner prescribed diagnosis details with the International Statistical Classification of Diseases and Related Health Problems (ICD).<\/p>\n<p>Diseases are classified with ICD-10 code; that code should have mapped with practitioner&#8217;s prescribed diagnosis.<\/p>\n<p>Here the challenges are customer&#8217;s data\u00a0are not aligned in right format and the diagnosis information varies\u00a0for each practitioner.<\/p>\n<p>Objective is, practitioner prescribed diagnosis information should\u00a0closely match\u00a0with International Statistical Classification of Diseases to map the ICD-10 code.<\/p>\n<h3 class=\"western\">Solution<\/h3>\n<p>Pentaho+ provides the solution for the above business case using PDI. With the help of Pentaho+ we can analyze the text file and process the semi-structured data to match with structured data.<\/p>\n<p>To increase match relevance we shall use the SOLR and OpenNLP techniques for the better search.<\/p>\n<h4>How PDI works for Fuzzy match<\/h4>\n<p>Pentaho+ data integration supports fuzzy match method. The Fuzzy Match step finds strings that potentially match using duplicate-detecting algorithms that calculate the similarity of two streams of data.<\/p>\n<p>This step returns matching values as a separated list as specified by user-defined minimal or maximal values.<\/p>\n<p>The below Algorithms are used in Pentaho Plus Fuzzy match step. Within the Algorithm field, there are several options available to compare and match strings.<\/p>\n<ul>\n<li><b>Levenshtein and Damerau-Levenshtein<\/b>&#8212;calculate the distance between two strings by looking at how many edit steps are needed to get from one string to another. The former only looks at inserts, deletes, and replacements. The latter adds transposition. The score indicates the minimum number of changes needed. For instance, the difference between John and Jan would be two; to turn the name John into Jan you need one step to replace the <i>O<\/i> with an <i>A<\/i>, and another step to delete the <i>H<\/i>.<\/li>\n<li><b>Needleman Wunsch<\/b>&#8212;calculates the similarity of two sequences and is mainly used in bioinformatics. The algorithm calculates a gap penalty. The aforementioned example would have a score of negative two.<\/li>\n<li><b>Jaro and Jaro Winkler<\/b>&#8212;calculate a similarity index between two strings. The result is a fraction between zero, indicating no similarity, and one, indicating an identical match.<\/li>\n<\/ul>\n<p style=\"padding-left: 30px\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-245\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog3-3.png\" alt=\"\" width=\"271\" height=\"183\" title=\"\"><\/p>\n<ul>\n<li><b>Pair letters similarity<\/b>&#8212;dissects the two strings in pairs and calculates the similarity of the two strings by dividing the number of common pairs by the sum of the pairs from both strings.<\/li>\n<\/ul>\n<p><b>Metaphone, Double Metaphone, SoundEx, and Refined SoundEx<\/b>&#8212;are phonetic algorithms, which try to match strings based on how they would sound. Each is based on the English language and would not be useful to compare other languages.<\/p>\n<ul>\n<li style=\"list-style-type: none\">\n<ul>\n<li>The Metaphone algorithm returns an encoded value based on the English pronunciation of a given word. The encoded value of the names John and Jan would return the value <i>JN<\/i> for both names.<\/li>\n<li>The Double Metaphone algorithm has fundamental design improvements over its predecessor and uses a more complex rule set for coding. It can return a primary and a secondary encoded value for a string. The names John and Jan each return Metaphone key values of <i>JN<\/i> and <i>AN<\/i>.<\/li>\n<li>The Soundex algorithm returns a single encoded value for a name that consists of a letter followed by three numerical digits. The letter is the first letter of the name, and the digits encode the remaining consonants.<\/li>\n<li>The Refined SoundEx algorithm is an improvement over its predecessor. Encoded values for this algorithm are six digits long, the initial character is encoded, and multiple possible encodings can be returned for a single name. Using this algorithm, the name John returns the values 160000 and 460000, as does the name Jan.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<h4>Process Flow in PDI:<\/h4>\n<p>To extract\/map the relevant data from the structured data for the semi-structured data, we used to follow the below steps<\/p>\n<p><strong>Step 1:<\/strong> wi Transform the semi structured data into structured data using respective transform steps from PDI<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-246\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog4-2.png\" alt=\"\" width=\"533\" height=\"320\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog4-2.png 533w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog4-2-300x180.png 300w\" sizes=\"auto, (max-width: 533px) 100vw, 533px\" \/><br \/>\n<b>Step 2<\/b>:\u00a0 Match the practitioner prescribed diagnosis with Master Diagnosis to fetch the respective ICD code<\/p>\n<p><b>Step 3<\/b>: Use the fuzzy match step to achieve the process, there we have to choose the right algorithm to do the fuzzy match process.<\/p>\n<p>Jaro Winkler algorithm will fetch maximum relevant strings from master with score value<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-247\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog5-1-1.png\" alt=\"\" width=\"682\" height=\"361\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog5-1-1.png 682w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog5-1-1-300x159.png 300w\" sizes=\"auto, (max-width: 682px) 100vw, 682px\" \/><br \/>\n<b>Step 4<\/b>: Store the mapped\/matched value into table with score value which is used to determine the matched sentence\u00a0accuracy<\/p>\n<p><b><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-248\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog6-1.png\" alt=\"\" width=\"404\" height=\"146\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog6-1.png 404w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog6-1-300x108.png 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog6-1-400x146.png 400w\" sizes=\"auto, (max-width: 404px) 100vw, 404px\" \/><br \/>\n<\/b><\/p>\n<h4>User Interface to Edit \/ Remap the ICD:<\/h4>\n<p>Pentaho Plus UI supports the customized table editor, which helps to update or remap the mapped the ICD code to the respective diagnosis.<\/p>\n<p>User can search the Diagnosis based on ICD code or ICD code based on diagnosis respectively.<\/p>\n<p>The below screenshot explains, how the interface can be used to update the matched records.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-249\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog7.jpg\" alt=\"\" width=\"945\" height=\"529\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog7.jpg 945w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog7-300x168.jpg 300w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog7-768x430.jpg 768w\" sizes=\"auto, (max-width: 945px) 100vw, 945px\" \/><\/p>\n<h1 class=\"western\">Summary<\/h1>\n<p>With the help of Pentaho Plus fuzzy match process, customer&#8217;s semi-structure data was analyzed and processed to match the standard classification of diseases.<\/p>\n<p>Customers are able to edit or remap ICD code for the respective diagnosis with the help of matching score value.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-250\" src=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog8-1.png\" alt=\"\" width=\"746\" height=\"412\" title=\"\" srcset=\"https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog8-1.png 746w, https:\/\/tenthplanet.in\/blogs\/wp-content\/uploads\/sites\/21\/2018\/02\/blog8-1-300x166.png 300w\" sizes=\"auto, (max-width: 746px) 100vw, 746px\" \/><\/p>\n<h3 class=\"western\">References:<\/h3>\n<p>Fuzzy match Step<\/p>\n<p><b>:https:\/\/wiki.pentaho.com\/display\/EAI\/Fuzzy+match<\/b><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Fuzzy matching is a method that provides an improved ability to process word-based matching queries to find matching phrases <\/p>\n","protected":false},"author":23,"featured_media":1131,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[424],"tags":[448,449,450],"class_list":["post-244","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-pentaho","tag-fuzzy-logic","tag-matching-words","tag-phonetically"],"acf":[],"_links":{"self":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts\/244","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/users\/23"}],"replies":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/comments?post=244"}],"version-history":[{"count":0,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/posts\/244\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/media\/1131"}],"wp:attachment":[{"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/media?parent=244"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/categories?post=244"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tenthplanet.in\/blogs\/wp-json\/wp\/v2\/tags?post=244"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}