Lift

 
Like its name indicates a lift is a measure on how good your binary classifier model is lifting the predictions. In other words it measures how well the new models choice is better than an old models choice or the random selection. A lift plot is an alternative to a ROC curve when you want to compare two classifier, it provides insights on the model and can help to determine a cutoff.

Lets start with the definition of list:
The lift measures the change in concentration of a target value if applied to a subgroup of the test set that is chosen by the model. Note that the lift is always connected to the concentration of the target value in the test set, the lower the concentration, the higher is also the possible lift of the model. Therefore there are no restrictions to the values list of the lift.

In an example:
Imagine that you want to improve a marketing campaign which adresses customers in order to sell new products or services. From the past campaign you build up a predictive model trying to identify the customers who are likely to respond. In the past campaign 4% of the overall adressed customer responded. If you now choose a random subset of 10 % of the adressed customers, you would expect 4% of the responders in there.
Now with your new model you can identify possible responders and you therefore choose the 10% most likely to respond to the campaign. If from these 4% of all customers 16% respond, then your classifier has a lift of 16 / 4 = 6.
Now you can calculate the expected responses adressing 20%, 30%, ... and you can plot the data in a so called lift chart, that could look like this (the yellow line corresponds to the random classifier, the red one to the fictious model):



These lift charts can be built up for different classifiers in order to compare them. Also from the form of the curve a possible maximum of adressed customer can be choosen in order to optimize the costs for the campaigns.





Classical Decomposition of a time series

A time series is a data set that has a time component. Yes, it is just what you think about, in the optimal case you have one value in a fixed time interval. It is a pretty actual topic to create massive datasets on behalf of sensors in an Internet of Things scenario, and you usually get too much values (>100 measurements per second), so to create a useful time series requires some preprocessing (filter or average). Missing values can be a problem, also it is usually not difficult to find a good estimate.

You can therefore build a time line which could look like this:


Here I used R's default dataset "AirPassengers" which reflects the monthly international airline passenger numbers in the years 1949 to 1960.

Natural questions are: 
  • Can you see which of the years were good, which of them were bad?
  • Can you split up the time series into more homogenious components?
  • Can you predict the values for future months? And how good are these predictions?
Time series analysis seems not too complex, however in reality often there are combinations of models required to make good predictions. In this post I want to show you how standard approches work. In reality, especially in retail, the timelines are usually not that easy to handle, as it is widely influenced by customer reviews.

The central idea is, that a time series $Y = (Y_t)$ is a combination of a three independent sub - time series:
- A trend component T is a long-term tendence in the data, it does not have to be linear.
- A seasonal component S is a pattern that reoccurs regularily after a fixed period (like every summer, every january or every day at 10:30).
- A random component I, also called irregular or noise.

We want to try to find these three time series in the upper mentioned example. First we have to decide on the type of decomposition, we can choose from additive and multiplicative.
In an additive model we add the 3 sub time series up to get the original time series: $$Y_t = T_t + S_t + I_t$$ You should use it when the seasonal variance does not change that much.

In a multiplicative model we multiply the 3 sub time series: $$Y_t= T_t * S_t * I_t$$ Use it when you see the peeks growing with time, like in the earlier mentioned example of airplane passengers. Here we should go for a multiplicative model.



Tip: A multiplicative model often can be changed into an additive model using the log function.

How would we get the values for the trend, seasonal and random sub time series? We will go step by step, just to motivate, here the result calculated by R with function decompose:


Here the corresponding coding in R: $$plot(decompose(ts(AirPassengers, frequency = 12, start =1949), type = "mult"))$$

To go on:

1. Here is how to determine the trend component
2. Here is how to determine the seasonal and random component
3. Here is a summary on the classical decomposition of time series

Seasonality and Random Determination

In this post we saw how the three components trend, seasonality and random of a time series. How to extract the trend was shown there, now we focus on how the seasonal component and the random component is determined.
Assume we have a detrended time series (we take here the AirPassengers time series and remove the trend). We assume a seasonality of a fixed period. In reality the assumption to have a fixed seasonality is too strict, as the period could shorten or change its structure over time. But under this assumption the determination of the seasonality is easy: To get the seasonal value of January, we take all values of January and build the average. This is the pattern we use for all periods.
The last step is to determine the random component $I$, we get it by simply removing the trend $T$ and seasonal component $S$ from the original time series $Y$, in an additive model this would be $I_t = Y_t - T_t - S_t$ and in a multiplicative model $I_t = Y_t / (T_t * S_t)$.
In our example, this is how the random components looks like:
What can we get out of it?
The random component shows the noise in the data, the values that do not fit the model. It helps to get a feeling how well the data is explained by the assumption to have a trend and a seasonality. The classical decomposition also could help to find outliers, which will show up with a high peek.

For completeness here again the whole picture holding all the steps discussed:



 

Trend determination with moving averages

We already saw in the previous post that you can decompose a trend from a time series. Here a classical approach how to determine an underlying trend.
 
A trend $T$ is actually a smoothened version a time series, it helps to capture global tendencies. To get a trend line here is what you have to do:
  •  Determine the seasonal periodicity of the time series (if there is one). These periodic patterns are usually visible, but if you cannot see them form the plotted chart, there are also methods using Fourier Transform Algorithms to determine them. In our example we see a yearly periodicity, as the values for the airplane passengers come in monthly, the periodic value is m=12.
  • With this number m use methods like the moving average of order m to determine the values of the trend

Now what is the moving average?
Once you understand the concept, it is easy to remember: Imagine that your dataset consists of 5 values $y_1$, ..., $y_5$. To determine the value of the trend of order $m = 3$ you would take the original value, the value of the predecessor and the value of the successor and create an average. In the simplest approach you simply take the sum of the values devided by the number of values (the so called Simple Moving Average SMA).
In the example of 5 values this would look like:
           
$y_1$      3
$y_2$      5     -> (3+5+4)/3 = 4
$y_3$      4     -> (5+4+1)/3 = 3.33
$y_4$      1     -> (4+1+3)/3 = 2.67
$y_5$      3



You already see the problem here: for the beginning and the ending there are no values available, the missing tail depends on $m$.

What if we choose $m=4$? We would have to decide to take more points from past than from future. In these cases the algorithm is not symmetric anymore, usually you therefore either change to the next odd number ($m=13$) or you choose a so called centered moving average. In the centered moving average you use a simple moving average of order 2 in the first step to determine values like $y_{1.5}$. Then you use the SMA of order $m$.

$y_1$      3
                           -> $y_{1.5}$ = (3+5)/2 = 4
$y_2$      5
                           -> $y_{2.5}$ = (5+4)/2 = 4.5
$y_3$      4                                                                           ->  (4+4.5+2.5+2) / 4 = 3.25
                           -> $y_{3.5}$ = (4+1)/2 = 2.5
$y_4$      1
                           -> $y_{4.5}$ = (1+3)/2 = 2
$y_5$      3

In our earlier example of air passengers we determined $m=12$ (even) and the trend values in yellow are determined by a centered moving average of order 12.

Here the command in R (ma stands for Moving Average):
lines(ma(x, order = 3, centre = TRUE))

After successfully determining the trend, we can remove it from the original data. In an additive model we get the de-trended time series by substracting it ($Y_t - T_t$), in a multiplicative model by deviding $Y_t/T_T$. This detrended time series of our AirPassenger example looks like this:
The next step is now to get the seasonality component and the random component from the de-trened time series.





Accuracy and Precision - How good is your classifier?


After successfully deciding on a classifier, after adjusting and optimizing the parameters on behalf of training data, it is time to evaluate the model. In the context of model evaluation, a few statistical terms are of importance, most importantly:

Accuracy and Precision

Both have colloquially pretty much the same meaning, however in statistics they describe different properties. They are independent of each other, models can be highly accurate but have low precisiony or also highly precision but low in accuracy. Ideally you have a model that is both highly accurate and highly pricise.

To understand the definitions, we start with the so called Confusion Matrix, which can be created for all classifiers or supervised learning models. After building up your model you compare the actual results to the results predicted by your model. You then build a matrix classifying the compared results via the following matrix (in case you have more results than true/false, you can create the confusion matrix to each feature):


The True Positives (TP) and the True Negatives (TN) are the datasets that your model correctly predicted, the higher the value in these boxes, the better is your model. Errors in the predicted classifications of your model can be devided in:
- False Positives (FP) (or also called Type-1-Errors) are the cases, in which your model classified a dataset incorrectly as positive,
- False Negatives (FN) (also called Type-2-Errors) are the ones incorrectly classified to negative.

Accuracy now is the ration of correctly classified datasets compared with the whole dataset, so
Accuracy = ( True Positives + True Negatives ) / All classified examples.

A high accuracy seems to be a reasonable choice to evaluate your model, however it is not sufficient, as the following example shows: Imagine you have a classifier for breast cancer, which in most cases will give negative result. In fact for 2017, US expect a rate of ~0,123% of new cases. Imagine that your classifier always predicts negative results, then your classifier will have TP = 0,987, TN = 0, and therefore an accuracy of 98,7%, however we agree that the classifier is not at all useful.
We somehow need to have to consider the overall positives (P) and the overall negatives (N).

Here is where precision comes into play. Precision is calculated as
Precision =  True Positives / ( True Positives + False Positives )

Now for our dummy classifier the precision would be 0 confirming that it is not useful at all.

There are more key figures relevant, but they will be posted on a later post. For now the confusion matrix is a first hint to determine if your classifier is useful.







How is the weather in the Black Forest? - Nearest Neighbour Approaches


Planning a city or hiking trip on the weekend, a barbequeue or a romantic picnic in the mountains requires a pretty good weather forecast. What would you do if you wanted to know the weather? Eventually you would just check the weather forecast in the desired city, these weather forcasts are availyble online. But what if your route is lying in the mountains, without any reference city or village? You would then choose a location near your route, for which a weather forecast is available and assume you will get the same weather on your route.
Assume you have a sunny weather forecast for one side of your route and rainy weather forecast on the other side, you would probably rely on the forecast closer to the route, but also consider the forcast farer away.
This approach is quite reasonable assuming that the weather is similar on geographic similar locations. The same reasoning is also applied in the business and research world, to forecast key figures like preferences, behaviours and prizes. One approach is the Nearest Neighbour Approach which starts with the assumption that objects "close" to each other behave in the "same" way, so if you have to predict an unknown value, just ask the "nearest" neighbours and use their results.

However there are some problems in this approach: it might be easy to find neighbours considering numeric values like age, income or location, but how would you find neighbours for, lets say, prefered music?

As often there is no general answer and different ways to find neighbours exists. Usually the tasks to find a good neighbour consist of two main tasks:
1. Find a measure for distance between two datasets
2. Combine the datasets for the closest neighbours to make a prediction

Usually you have a dataset consisting of a collection of features, for a customer this could be a so called Customer ID hodling all the relevant data that could possibly influence the target variable. Often age, income, location are part of it. For these features is it easy to determine a distance, usually the difference (direct difference or euklidian difference) is used. For better comparison normalization is recommended, as the features have different ranges.

You could then go for the k nearest neigbours, check the known values of the target variables and would the have to combine these results to a target value. Here business knowledge is required to determine to which extend e.g. the age is important or if location is more relevant than income.

The advantages of this approach is that they are usually pretty easy to understand, often show additional insights and change whenever the data changes.
However they could be time consuming as for finding good neighbours you would have to check every single known dataset and measure the distance. Also the prediction could depend considerably on the value of k, further analysis is required. And the algorithm is discontinuous, meaning that a new dataset could have a huge impact on the existing predictions.

To overcome the bounderies different approaches have been established (e.g. reducing the number of datasets by choosing "important" neigbours), and the nearest neigbour approaches ars widely used in different classification or regression problems (e.g. supporting breast cancer detection, estimations on house prizes).
And defining the right distance measure you can even find neighbours to music files or detect song titles.

As a closing example here some simplified excercise. Consider the following data:

Target: Money spent online for hobbies per year
Gender Age    Income Nof Children Target
Male     35      65.000            1            6.500
Female 24      25.000           0           2.000
Female   60      60.000            4              380
Male      48      45.000            2           2.100
Female   39     60.000            0            6.000
Male      49       75.000            2  7.500
Male      18           800             0               670

Which are your closest neigbours? And how good does the estimation of the 1-nearest neighbour model fit to you?
For the distance choose the absolute difference value in age, income and number of children, take the weights age_weight = 7, income_weight = 10, nof_children_weight = 1.
(Note that this is a very simplified example, the money spent on hobbies would in reality rely on way more factors and the variables are not independent of each other).

Hopefully you get a good approximation, so your next romantic outdoor picnic in the Black Forest does not look like this:

What does a Data Scientist do? - Facts



From the post about the CRISP-DM we know what phases a data science project consist of. This post analyzes the time spend on the different tasks. I refer to the "Data Science Survey" of Rexer Analytics from 2015.
  • The top application areas are Marketing, Academics, Finance, Technology, Medical, Retail, Internet-based, Government and Manufacturing.
  • Data scientists work above all as corporate employees, consultants, academics or vendors.
  • Data scientists have a high job satisfaction level.
  • The biggest challenges are the rising data analysis complexity and data visualization.
  • The most frequently used algorithms are regression algorithms, decision trees and cluster analysis, the main tool is R.
  • There is also an impressive list of alternative job titles provided, so apart from data scientists, you can also call yourself data analyst, researcher, business analyst, data miner, statistican, predictive modeler, computer scientist, engineer, software developer.

Data Scientists - Skills


It is obvious that a data scientist should have an interest in data, a strong analytical, mathematical background and a good intuition on numbers. A profund basic of statistics (the more the better), an insight into the tools (Excel is a very powerful and useful tool!) and a good part of endurance to handle complex processes and work on large data sets.
However these skills are not the essential competencies of a data scientist, moreover it is important that a data scientist understands the details on the algorithms applied in order to decide how to further optimize models, performance and processes. Experience with performance optimization algorithms, distributed calculation and programming are of great help. A good data scientist is able to structure and prioritize, distinguish between useful and unuseful information and also be able to work with people from different departments. Further imprescindible abilities are good communication and presentation skills to a wide audience, the task to help business in decisions on complex business processes is only fulfilled if the business actually understands and gets to "see" the results.
Last but not least a data scientist should always be open to new ideas, as the models and the requirements change regularily.

The Data Mining Process


The Cross Industry Standard Process for Data Mining (CRISP-DM) is a process model that describes the different steps data scientists use for analyzing data. Here some more details on the different steps:

In the Business Understanding Phase, you describe the current situation and the requirements from a business point of view, here the business defines, what it wants to accomplish with the data mining process. This included not only a business project plan, a list of risks and measures, costs and benefits, limitations and terminologies, but also a business success criteria.
In the Data Understanding Phase, the business requirements are translated into a data science problem. Here you describe the problem in technical terms, define inputs and outputs, detect dependencies and create a dynamic technical project plan. Here you also describe the data mining success criteria in a simple and easily understandable way. This coulds be a goal like the creation of BI-reports to learn from the historical data, cluster models, segmentations for unsupervized/undirected learning or a prediction, risk estimation or a classification for supervised/directed learning.
In both upper mentioned steps, the requirements should be described as concrete as possible.

The next phase is the Data Preparation Phase, in which an initial data set is being collected from documented tables in documented locations, described (missing fields, amounts of data, relations,...), explored (first findings, simple associations, initial hypothesis,...)  and verified (is the collected data correct/useful/intuitive?). Here you also transform class categories into numerical values, aggregate suitable datasets, clean and format data (eg. unsort data for neural networks).

The Modeling Phase stands in close contact to the Data Preparation Phase. In it the available data is cut into training, test and cross validation set. Here starting with the initially retrieved data you define new records (e.g. areas from witdh and length), datasets from different tables are collected to create an analytical record (or DNA) of a target object (also called entity, e.g. a customer), that hold the whole amount of data needed. Here you also identify outliers, select and sample data to be used in a model and analyze groupings of data. You decide on the model to use and the features (delete unrelevant features, include new features).

In the Evaluation Phase you analyze the results of the model, the accurancy and robustness of the model, the error and optimization opportunities. With these results a discussion round with the business is needed in order to verify if the results from the evaluated data are meeting the business requirements and that all important insights and goals have been included into the analysis. Usually more than one round is needed.

In the Deployment Phase the model is deployed, so that the functionality can be handed over and taken care of by the business. This could be in form of a BI-report or in the implementation of programs that allow the regular update of data.

As the models are created for a certain point in time, a certain business goal and on historic data, the data mining process has to be adapted and updated with new data.