For Wordage, we need to import the library from the gen sim.Here are some points sentences: the list of split sentences. size: the dimensionality of the embedding vector window the number of context words you are looking at mincount: tells the model to ignore words wit
Data Science
Answered Dec 24, 2019
·
1.1K Views
Read answer →
First we import the data and convert the target feature as factor # Importing the dataset dataset = read.csv('Social_Network_Ads.csv') dataset = dataset[3:5] Now we split, scale and fit the data into a KNN model # Splitting the dataset into the Training set and Test set # install.packages('caTools') library(caTools)…
Data Science
Answered Dec 20, 2019
·
1.4K Views
Read answer →
A data frame is a 2 dimensional structure containing data in the form of rows and columns. A data frame is used to easily represent data in the form of spreadsheet and can be easily accessed in every algorithm. Also a dataframe contains a series of dictionaries with particular keys and values that are represented in…
Data Science
Answered Dec 10, 2019
·
1.4K Views
Read answer →
#Below are the variables of my data. train_data.dtypes OUTPUT TripType category VisitNumber category Weekday category Upc category ScanCount int64
Data Science
Answered Nov 30, 2019
·
1.4K Views
Read answer →
ROC curve can give us a clear idea to set a threshold value to classify the label and also help in model optimization. A low threshold value we will put most of the predicted observations under the positive category, even when some of them should be placed under the negative category. On the other hand, keeping the…
Data Science
Answered Nov 30, 2019
·
1.1K Views
Read answer →
Maximum likelihood estimation is a method of estimating the parameters of a model given observations, by finding the parameter values that maximize the likelihood of making the observations, this means finding parameters that maximize the probability p of event 1 and (1-p) of non-event 0, as we know: probability…
Data Science
Answered Nov 28, 2019
·
1.3K Views
Read answer →
To remove non alphanumeric symbols from text data, regex is used. A regex basically used .compile() method to remove all symbols and sub() method to separate all the words with spaces. Suppose a text contains the following A B C D ,5 .. AAA55AAA aaa.bbb.ccc And the output should be 'A' 'B' 'C' 'D' 'AAA' 'AAA' 'aaa'…
Data Science
Answered Nov 12, 2019
·
1.4K Views
Read answer →
First we import the data # Importing the libraries import numpy as np import matplotlib.pyplot as plt import pandas as pd # Importing the dataset dataset = pd.read_csv('Position_Salaries.csv') X = dataset.iloc[:, 1:2].values y = dataset.iloc[:, 2].values Then we split the data # Splitting the dataset into the Training…
Data Science
Answered Nov 9, 2019
·
1.4K Views
Read answer →
It is an impurity which is also a measure of misclassification which applies in a context of multi class classifier. It works similar to entropy but is quicker and easy to calculate. It works on the blow formula.Where i=number of classes. The similarity between Gini and entropy is shown below
Data Science
Answered Nov 7, 2019
·
1.1K Views
Read answer →
import pandas as pd from sklearn.tree import DecisionTreeClassifier data = pd.DataFrame() data['A'] = ['a','a','b','a'] data['B'] = ['b','b','a','b'] data['C'] = [0, 0, 1, 0] data['Class'] = ['n','n','y','n'] tree =
Data Science
Answered Nov 4, 2019
·
1.0K Views
Read answer →