- Decision Tree In Data Mining – An Important Guide For Beginners In 2021
- 1) What is the decision tree in data mining?
- 2) Decision Tree Algorithm in Data Mining
- 3) Important Terms of Decision Tree in Data Mining
- 4) Root nodes
- 5) Application of Decision Tree in Data Mining
- 6) Advantages of Decision Tree
- 7) Disadvantages of Decision Tree
- Conclusion
- ALSO READ
- Методы классификации и прогнозирования. Деревья решений
Decision Tree In Data Mining – An Important Guide For Beginners In 2021
1) What is the decision tree in data mining?
A decision tree is a plan that includes a root node, branches, and leaf nodes. Every internal node characterizes an examination on an attribute, each division characterizes the consequence of an examination, and each leaf node grasps a class tag. The primary node in the tree is the root node.
The subsequent decision tree is for the thought buy a computer that shows whether a purchaser at an enterprise is expected to buy a computer or not. Each internal node characterizes an inspection on an attribute. Each leaf node signifies a class.
2) Decision Tree Algorithm in Data Mining
Decision Tree algorithm relates to the persons of directed intelligence techniques. Unlike other-directed education procedures, the decision tree algorithm can be used to answer deterioration and arrangement difficulties.
The objective of using a Decision Tree is to craft a preparation ideal that can use to foresee the class or value of the mark variable by learning easy judgement procedures incidental from previous information (training data).
In Decision Trees, for estimating a class tag for best ever we start with the root of the tree. We make relations with the root attribute to the record’s attribute. We make division agreeing to that value and jump to the subsequent node on the base of choice.
3) Important Terms of Decision Tree in Data Mining
Decision trees can handle complicated data, which is a portion of what results in them valuable. Though, this doesn’t mean that they are difficult to know. At their centre, all decision trees finally include three vital portions or nodes.
- Decision nodes: Represents a decision and is normally displayed with a square.
- Chance nodes: Represents chance or confusion and is normally displayed with a circle.
- End nodes: Represents a result and is normally displayed with a triangle.
By connecting these different nodes, we get divisions. We can use nodes and divisions an unlimited number of times to form trees of different difficulties. Let’s see how these portions appear before we include any information.
Fortunately, many decision tree vocabulary keep an eye on the tree equivalence, which marks it full calmer to recollect! Let’s explore these terminologies now:-
4) Root nodes
The blue decision is called the ‘root node’. This is at all the times the primary node in the path. It is the knot from which all other choices, forecasts and end knots finally divide.
In the figure above, the lavender end nodes are called the ‘leaf nodes.’ These display the conclusion of a decision route (or outcome). Every time you recognize a leaf node because it doesn’t fragment, or subdivide any more like a real leaf.
In between the origin knots and the leaf knots, we can have any number of internal ties. These can comprise decisions and chance nodes (for ease, this image only uses chance nodes). It is really easy to identify an internal node as each internal nodes have branches of its own while also joining to the earlier node.
Dividing or ‘splitting’ is said when any node divides two or more substitute nodes. These substitute nodes can also be another internal node, or they can tip to result (a leaf/ end node)
Rarely decision trees can become attractively miscellaneous. In these circumstances, they can close up giving too much load to immaterial information. To sidestep this difficulty, we can eliminate definite nodes using a procedure well known as ‘pruning’. Pruning is precisely what it echoes like if the tree develops branches we don’t require, we basically cut them off.
5) Application of Decision Tree in Data Mining
Notwithstanding their disadvantages, decision trees are static an influential and prevalent means. They are usually used by information experts to bring out an analytical investigation (e.g., improve procedures policies in trades). They are to a prevalent means for machine learning and artificial intelligence, where they are used as preparation procedures for administered wisdom (i.e. classifying information based on various tests, such as ‘sure’ or ‘nope’ classifiers.)
Mostly, decision trees are used in an extensive variety of businesses, to crack numerous categories of difficulties. Because of their elasticity, they are used in areas from know-how and fitness to the fiscal formation. Illustrations comprise:
- A know-how corporate assessing extension occasions based on examination of earlier revenue information.
- A puppet business determining where to objective its partial marketing financial strategy, based on what demographic information guides consumers is likely to purchase.
- Banks and loan providers using past information to forecast how likely it is that a debtor will default on their payments.
6) Advantages of Decision Tree
- In comparison to other procedures, decision trees need not as much energy for information training during pre-processing.
- A decision tree does not involve stabilization of information.
- A decision tree does not need scaling of information as well.
- Omitted values in the information also do not disturb the procedure of constructing a decision tree to any substantial degree.
- A Decision tree model is identical instinctive and stress-free to describe to practical squads as well as investors.
7) Disadvantages of Decision Tree
- A minor variation in the information can cause a huge variation in the configuration of the decision tree triggering unpredictability.
- For a Decision tree occasionally calculation can go far extra multifaceted in comparison to other procedures.
- Decision tree repeatedly takes greater time to train the model.
- Decision tree preparation is comparatively lavish as the difficulty and period have taken are extra.
- The Decision Tree procedure is insufficient for relating deterioration and forecasting uninterrupted values.
Conclusion
Decision Trees helps to forecast upcoming events and are easy to understand. They work more efficiently with discrete attributes. They may suffer from error propagation.
If you are interested in making a career in the Data Science domain, our 11-month in-person Postgraduate Certificate Diploma in Data Science course can help you immensely in becoming a successful Data Science professional.
ALSO READ
decision tree algorithm in data mining decision tree explained decision tree in data mining decision tree in data mining example decision tree induction in data mining what is decision tree in data mining
Источник
Методы классификации и прогнозирования. Деревья решений
Аннотация: Описывается метод деревьев решений. Рассматриваются элементы дерева решения, процесс его построения. Приведены примеры деревьев, решающих задачу классификации. Даны алгоритмы конструирования деревьев решений CART и C4.5.
Метод деревьев решений ( decision trees ) является одним из наиболее популярных методов решения задач классификации и прогнозирования. Иногда этот метод Data Mining также называют деревьями решающих правил , деревьями классификации и регрессии.
Как видно из последнего названия, при помощи данного метода решаются задачи классификации и прогнозирования.
Если зависимая, т.е. целевая переменная принимает дискретные значения, при помощи метода дерева решений решается задача классификации.
Если же зависимая переменная принимает непрерывные значения, то дерево решений устанавливает зависимость этой переменной от независимых переменных, т.е. решает задачу численного прогнозирования.
Впервые деревья решений были предложены Ховилендом и Хантом (Hoveland, Hunt) в конце 50-х годов прошлого века. Самая ранняя и известная работа Ханта и др., в которой излагается суть деревьев решений — «Эксперименты в индукции» («Experiments in Induction «) — была опубликована в 1966 году.
В наиболее простом виде дерево решений — это способ представления правил в иерархической, последовательной структуре. Основа такой структуры — ответы «Да» или «Нет» на ряд вопросов.
На рис. 9.1 приведен пример дерева решений, задача которого — ответить на вопрос: «Играть ли в гольф?» Чтобы решить задачу, т.е. принять решение, играть ли в гольф, следует отнести текущую ситуацию к одному из известных классов (в данном случае — «играть» или «не играть»). Для этого требуется ответить на ряд вопросов, которые находятся в узлах этого дерева, начиная с его корня.
Первый узел нашего дерева «Солнечно?» является узлом проверки , т.е. условием. При положительном ответе на вопрос осуществляется переход к левой части дерева, называемой левой ветвью , при отрицательном — к правой части дерева. Таким образом, внутренний узел дерева является узлом проверки определенного условия. Далее идет следующий вопрос и т.д., пока не будет достигнут конечный узел дерева, являющийся узлом решения . Для нашего дерева существует два типа конечного узла : «играть» и «не играть» в гольф.
В результате прохождения от корня дерева (иногда называемого корневой вершиной) до его вершины решается задача классификации, т.е. выбирается один из классов — «играть» и «не играть» в гольф.
Целью построения дерева решения в нашем случае является определение значения категориальной зависимой переменной.
Итак, для нашей задачи основными элементами дерева решений являются:
Внутренний узел дерева или узел проверки : «Температура воздуха высокая?», «Идет ли дождь?»
Лист , конечный узел дерева, узел решения или вершина : «Играть», «Не играть»
Ветвь дерева (случаи ответа): «Да», «Нет».
В рассмотренном примере решается задача бинарной классификации , т.е. создается дихотомическая классификационная модель. Пример демонстрирует работу так называемых бинарных деревьев.
В узлах бинарных деревьев ветвление может вестись только в двух направлениях, т.е. существует возможность только двух ответов на поставленный вопрос («да» и «нет»).
Бинарные деревья являются самым простым, частным случаем деревьев решений. В остальных случаях, ответов и, соответственно, ветвей дерева, выходящих из его внутреннего узла , может быть больше двух.
Рассмотрим более сложный пример. База данных , на основе которой должно осуществляться прогнозирование, содержит следующие ретроспективные данные о клиентах банка, являющиеся ее атрибутами: возраст, наличие недвижимости, образование, среднемесячный доход, вернул ли клиент вовремя кредит . Задача состоит в том, чтобы на основании перечисленных выше данных (кроме последнего атрибута) определить, стоит ли выдавать кредит новому клиенту.
Как мы уже рассматривали в лекции, посвященной задаче классификации, такая задача решается в два этапа: построение классификационной модели и ее использование.
На этапе построения модели, собственно, и строится дерево классификации или создается набор неких правил . На этапе использования модели построенное дерево , или путь от его корня к одной из вершин, являющийся набором правил для конкретного клиента, используется для ответа на поставленный вопрос «Выдавать ли кредит ?»
Правилом является логическая конструкция, представленная в виде «если : то :».
На рис. 9.2. приведен пример дерева классификации, с помощью которого решается задача «Выдавать ли кредит клиенту?». Она является типичной задачей классификации, и при помощи деревьев решений получают достаточно хорошие варианты ее решения.
Как мы видим, внутренние узлы дерева (возраст, наличие недвижимости, доход и образование) являются атрибутами описанной выше базы данных . Эти атрибуты называют прогнозирующими, или атрибутами расщепления (splitting attribute ). Конечные узлы дерева, или листы, именуются метками класса, являющимися значениями зависимой категориальной переменной «выдавать» или «не выдавать» кредит .
Каждая ветвь дерева, идущая от внутреннего узла , отмечена предикатом расщепления . Последний может относиться лишь к одному атрибуту расщепления данного узла. Характерная особенность предикатов расщепления : каждая запись использует уникальный путь от корня дерева только к одному узлу-решению. Объединенная информация об атрибутах расщепления и предикатах расщепления в узле называется критерием расщепления (splitting criterion ) [33].
На рис. 9.2. изображено одно из возможных деревьев решений для рассматриваемой базы данных . Например, критерий расщепления «Какое образование?», мог бы иметь два предиката расщепления и выглядеть иначе: образование «высшее» и «не высшее». Тогда дерево решений имело бы другой вид.
Таким образом, для данной задачи (как и для любой другой) может быть построено множество деревьев решений различного качества, с различной прогнозирующей точностью.
Качество построенного дерева решения весьма зависит от правильного выбора критерия расщепления . Над разработкой и усовершенствованием критериев работают многие исследователи.
Метод деревьев решений часто называют «наивным» подходом [34]. Но благодаря целому ряду преимуществ, данный метод является одним из наиболее популярных для решения задач классификации.
Источник