# Notebook 5 - pandas
[pandas](http://pandas.pydata.org) provides high-level data structures and functions designed to make working with structured or tabular data fast, easy and expressive. The primary objects in pandas that we will be using are the `DataFrame`, a tabular, column-oriented data structure with both row and column labels, and the `Series`, a one-dimensional labeled array object.

pandas blends the high-performance, array-computing ideas of NumPy with the flexible data manipulation capabilities of spreadsheets and relational databases. It provides sophisticated indexing functionality to make it easy to reshape, slice and perform aggregations.

While pandas adopts many coding idioms from NumPy, the most significant difference is that pandas is designed for working with tabular or heterogeneous data. NumPy, by contrast, is best suited for working with homogeneous numerical array data.
<br>

## Table of Contents:
- [Data Structures](#structures)
    - [Series](#series)
    - [DataFrame](#dataframe)
- [Essential Functionality](#ess_func)
    - [Reindexing](#reindexing)
    - [Dropping Entries](#removing)
    - [Indexing, Slicing and Filtering](#indexing)
    - [Arithmetic Operations](#arithmetic)
- [Summarizing and Computing Descriptive Statistics](#sums)
- [Loading and storing data](#loading)
    - [Text Format](#text) 
    - [Web Scraping](#web)
- [Data Cleaning and preperation](#cleaning)
    - [Handling missing data](#missing)
    - [Data transformation](#transformation)

The common pandas import statment is shown below:

In [1]:
# Common pandas import statement
import numpy as np
import pandas as pd

# Data Structures <a name="structures"></a>
## Series <a name="series"></a>
A Series is a one-dimensional array-like object containing a sequence of values and an associated array of data labels called its index.

The easiest way to make a Series is from an array of data:

In [2]:
data = pd.Series([4, 7, -5, 3])

Now try printing out data

The string representation of a Series displayed interactively shows the index on the left and the values on the right. Because we didn't specify an index, the default on is simply integers 0 through N-1.

You can output only the values of a Series using 
```python
data.values
```
or you can get only the indices using
```python
data.index
```
Try it out below!

You can specify custom indeces when intialising the Series

In [3]:
data2 = pd.Series([4, 7, -5, 3], index=["a", "b", "c", "d"])

Now you can use these labels to access the data similar to a normal array

In [4]:
data2["a"]

4

Another way to think about Series is as a fixed-length ordered dictionary. Furthermore, you can actually define a Series in a similar manner to a dictionary

In [5]:
cities = {"Glasgow" : 599650, "Edinburgh" : 464990, "Abardeen" : 196670, "Dundee" : 147710}
data3 = pd.Series(cities)

In [6]:
data3

Glasgow      599650
Edinburgh    464990
Abardeen     196670
Dundee       147710
dtype: int64

You can do arithmetic operations between Series similar to NumPy arrays. Even if you have 2 datasets with different data, arithmetic operations will be aligned according to their indices.

Let's look at an example

In [7]:
cities_uk = {"Birmingham" : 1092330, "Leeds": 751485, "Glasgow" : 599650,
             "Manchester" : 503127, "Edinburgh" : 464990}
data4 = pd.Series(cities_uk)

In [8]:
data3 + data4

Abardeen            NaN
Birmingham          NaN
Dundee              NaN
Edinburgh      929980.0
Glasgow       1199300.0
Leeds               NaN
Manchester          NaN
dtype: float64

Notice how some of the results are NaN? Well, that is because there were no instances of those cities within both of the datasets. You can usually extract NaNs from a Series with
```python
data4.isnull()
```

## DataFrame <a name="dataframe"></a>
A DataFrame represents a rectangular table of data and contains an ordered collection of columns, each of which can be a different value type. The DataFrame has both row and column index and can be thought of as a dict of Series all sharing the same index.

The most common way to create a DataFrame is with dicts

In [9]:
data = {"cities" : ["Glasgow", "Edinburgh", "Abardeen", "Dundee"],
        "population" : [599650, 464990, 196670, 147710],
        "year" : [2011, 2013, 2013, 2013]}
frame = pd.DataFrame(data)

Try printing it out

In [10]:
frame

Unnamed: 0,cities,population,year
0,Glasgow,599650,2011
1,Edinburgh,464990,2013
2,Abardeen,196670,2013
3,Dundee,147710,2013


Jupyter Notebooks prints it out in a nice table but the basic version of this is also just as readable!

Additionally you can also specify the order of columns during initialisation

In [11]:
frame2 = pd.DataFrame(data, columns=["year", "cities", "population"])

You can retrieve a particular column from a DataFrame with
```python
frame["cities"]
```
The result is going to be a Series

Additionally, you can retrieve a row from the dataset using
```python
frame[1]
```
Try it out below

It is also possible to add and modify the columns of a DataFrame

In [12]:
frame2["size"] = 100

In [13]:
frame2

Unnamed: 0,year,cities,population,size
0,2011,Glasgow,599650,100
1,2013,Edinburgh,464990,100
2,2013,Abardeen,196670,100
3,2013,Dundee,147710,100


In [14]:
frame2["size"] = [175, 264, 65.1, 60]  # in km^2

Similar to dicts, columns can be deleted using
```python
del frame2["size"]
```

Another common way of creating DataFrames is from a nested dict of dicts:

In [15]:
data2 = {"Glasgow": {2011: 599650},
        "Edinburgh": {2013:464990},
        "Abardeen": {2013: 196670}}

frame3 = pd.DataFrame(data)
frame3

Unnamed: 0,cities,population,year
0,Glasgow,599650,2011
1,Edinburgh,464990,2013
2,Abardeen,196670,2013
3,Dundee,147710,2013


Here is a table of different ways of initialising a DataFrame for your reference

| Type | Notes |
| --- | --- |
| 2D ndarray | A matrix of data; passing optional row and column labels |
| dict of arrays, lists, or tuples | Each sequence becomes a column in the DataFrame; all sequences must be the same length |
| dict of Series | Each value becomes a column; indexes from each Series are unioned together to<br>form the result's row index if not explicit index is passed |
| dict of dicts | Each inner dict becomes a column; keys are unioned to form the row<br>index as in the "dict of Series" case |
| List of dicts or Series | Each item becomes a row in the DataFrame; union of dict keys or<br>Series indices becomes the DataFrame's column labels |
| List of lists or tuples | Treated as the "2D ndarray" case |

# Essential Functionality <a name="ess_func"></a>
In this section, we will go through the fundamental mechanics of interacting with the data contained in a Series or DaraFrame.

## Reindexing <a name="reindexing"></a>
With pandas it is easy to restructure the order of your columns and rows using the `reindex` function. Let's have a look at an example:

In [16]:
# first define a new Series
s = pd.Series([1, 2, 3, 4, 5], index=['a', 'b', 'c', 'd', 'e'])
s

a    1
b    2
c    3
d    4
e    5
dtype: int64

In [17]:
# Now you can reshuffle the indices
s = s.reindex(['d', 'b', 'a', 'c', 'e'])
s

d    4
b    2
a    1
c    3
e    5
dtype: int64

Easy as that! This can also be extended for DataFrames, where you can reorder both the columns and indices at the same time!

In [18]:
# first define a new Dataframe
data = np.reshape(np.arange(9), (3,3))
df = pd.DataFrame(data, index=["a", "b", "c"],
                  columns=["Edinburgh", "Glasgow", "Aberdeen"])
df

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5
c,6,7,8


In [19]:
# Now we can restructure it with reindex
df = df.reindex(index=["a", "d", "c", "b"],
          columns=["Aberdeen", "Glasgow", "Edinburgh", "Dundee"])

Notice something interesting? We can actually add new indices and columns using the `reindex` method. This results in the new slots in our table to be filled in with `NaN` values.

## Removing columns/indices <a name="removing"></a>
Similar to above, it is easy to remove entries. This is done with the `drop()` method and can be applied to both columns and indices:

In [20]:
# define new DataFrame
data = np.reshape(np.arange(9), (3,3))
df = pd.DataFrame(data, index=["a", "b", "c"],
                  columns=["Edinburgh", "Glasgow", "Aberdeen"])

df.drop("b")

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
c,6,7,8


In [21]:
# You can also drop from a column
df.drop(["Aberdeen", "Edinburgh"], axis="columns")

Unnamed: 0,Glasgow
a,1
b,4
c,7


## Indexing, slicing and filtering <a name="indexing"></a>

### Indexing

Series indexing works analogously to NumPy array indexing (i.e. `data[...]`). You can also use the Series' index values instead of only integers:

In [22]:
s = pd.Series(np.arange(4), index=['a', 'b', 'c', 'd'])
s

a    0
b    1
c    2
d    3
dtype: int64

In [23]:
s[1]

1

In [24]:
s[3]

3

In [25]:
s["c"]

2

In [26]:
s[[1,3]]

b    1
d    3
dtype: int64

In [27]:
s[s<2]

a    0
b    1
dtype: int64

A subtle difference when indexing in pandas is that unlike in normal Python, slicing here is inclusive at the end-point.

In [28]:
s["b":"c"]

b    1
c    2
dtype: int64

All of the above also apply to DataFrames:

In [29]:
data = np.reshape(np.arange(9), (3,3))
df = pd.DataFrame(data, index=["a", "b", "c"],
                  columns=["Edinburgh", "Glasgow", "Aberdeen"])

df

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5
c,6,7,8


In [30]:
df[:2]

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5


In [31]:
df["Glasgow"]

a    1
b    4
c    7
Name: Glasgow, dtype: int64

### loc and iloc
For DataFrame label-indexing on the rows, you can use `loc` for labels and `iloc` for integer-indexing.

In [32]:
df.loc["b"]

Edinburgh    3
Glasgow      4
Aberdeen     5
Name: b, dtype: int64

In [33]:
df.loc["b", ["Glasgow", "Aberdeen"]]

Glasgow     4
Aberdeen    5
Name: b, dtype: int64

Now let's try `iloc`

In [34]:
df.iloc[1]

Edinburgh    3
Glasgow      4
Aberdeen     5
Name: b, dtype: int64

In [35]:
df.iloc[1, [1,2]]

Glasgow     4
Aberdeen    5
Name: b, dtype: int64

In [36]:
df.iloc[:2]

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5


Summary of indexing:

| Type | Notes |
| -- | -- |
| df\[val\]              | Select single column or sequency of columns from a DataFrame |
| df.loc\[val\]          | Select single row or subset of rows from a DataFrame by label |
| df.loc\[:, val\]       | Select single column or subset of columns by label |
| df.loc\[val1, val2\]   | Select both rows and columns by label |
| df.iloc\[idx\]         | Select single row or subset of rows from DataFrame by integer position |
| df.iloc\[:, idx\]      | Select single column or subset of columns by integer position |
| df.iloc\[idx1, idx2\]  | Select both rows and columns by integer position |
| reindex method         | Select either rows or columns by labels |

### Exercise 1
A dataset of random numbers is created below. Index the 47th column and the 22nd row. You should get the number **4621**.

*Note: Remember that Python uses 0-based indexing*

In [38]:
df = pd.DataFrame(np.reshape(np.arange(10000), (100,100)))

df.iloc[46,21]

4621

### Exercise 2
Using the same DataFrame from the previous exercise, obtain all rows starting from row 85 to 97.

In [40]:
df.iloc[84:96]

Unnamed: 0,0,1,2,3,4,5,6,7,8,9,...,90,91,92,93,94,95,96,97,98,99
84,8400,8401,8402,8403,8404,8405,8406,8407,8408,8409,...,8490,8491,8492,8493,8494,8495,8496,8497,8498,8499
85,8500,8501,8502,8503,8504,8505,8506,8507,8508,8509,...,8590,8591,8592,8593,8594,8595,8596,8597,8598,8599
86,8600,8601,8602,8603,8604,8605,8606,8607,8608,8609,...,8690,8691,8692,8693,8694,8695,8696,8697,8698,8699
87,8700,8701,8702,8703,8704,8705,8706,8707,8708,8709,...,8790,8791,8792,8793,8794,8795,8796,8797,8798,8799
88,8800,8801,8802,8803,8804,8805,8806,8807,8808,8809,...,8890,8891,8892,8893,8894,8895,8896,8897,8898,8899
89,8900,8901,8902,8903,8904,8905,8906,8907,8908,8909,...,8990,8991,8992,8993,8994,8995,8996,8997,8998,8999
90,9000,9001,9002,9003,9004,9005,9006,9007,9008,9009,...,9090,9091,9092,9093,9094,9095,9096,9097,9098,9099
91,9100,9101,9102,9103,9104,9105,9106,9107,9108,9109,...,9190,9191,9192,9193,9194,9195,9196,9197,9198,9199
92,9200,9201,9202,9203,9204,9205,9206,9207,9208,9209,...,9290,9291,9292,9293,9294,9295,9296,9297,9298,9299
93,9300,9301,9302,9303,9304,9305,9306,9307,9308,9309,...,9390,9391,9392,9393,9394,9395,9396,9397,9398,9399


## Arithmetic <a name="arithmetic"></a>
When you are performing arithmetic operations between two objects, if any index pairs are not the same, the respective index in the result will be the union of the index pair. Let's have a look

In [41]:
s1 = pd.Series(np.arange(5), index=["a", "b", "c", "d", "e"])
s2 = pd.Series(np.arange(4), index=["b", "c", "d", "k"])
s1 + s2

a    NaN
b    1.0
c    3.0
d    5.0
e    NaN
k    NaN
dtype: float64

The internal data alignment introduces missing values in the label locations that don't overlap. It is similar for DataFrames:

In [42]:
df1 = pd.DataFrame(np.arange(12).reshape((3,4)),
                  columns=list("abcd"))
df1

Unnamed: 0,a,b,c,d
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11


In [43]:
df2 = pd.DataFrame(np.arange(16).reshape((4,4)),
                  columns=list("cdef"))
df2

Unnamed: 0,c,d,e,f
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11
3,12,13,14,15


In [44]:
# adding the two
df1+df2

Unnamed: 0,a,b,c,d,e,f
0,,,2.0,4.0,,
1,,,10.0,12.0,,
2,,,18.0,20.0,,
3,,,,,,


Notice how where we don't have matching values from `df1` and `df2` the output of the addition operation is `NaN` since there are no two numbers to add.

Well, we can "fix" that by filling in the `NaN` values. This effectively tells pandas where there are no two values to add, assume that the missing value is just zero.

In [45]:
df1.add(df2, fill_value=0)

Unnamed: 0,a,b,c,d,e,f
0,0.0,1.0,2.0,4.0,2.0,3.0
1,4.0,5.0,10.0,12.0,6.0,7.0
2,8.0,9.0,18.0,20.0,10.0,11.0
3,,,12.0,13.0,14.0,15.0


Another important point here is although the normal arithmetic operations work here, there also exist dedicated methods like `DataFrame.add()` which achieve the same functionality + a bit extra.

Here's a list of all arithmetic operations within pandas:

| Operator | Method | Description |
| -- | -- | -- |
| + | add, radd | Addition |
| - | sub, rsub | Subtraction |
| / | div, rdiv | Division |
| // | floordiv, rfloordiv | Floor division |
| * | mul, rmul | Multiplication |
| ** | pow, rpow | Exponentiation |

Notice how some of the methods have `r` in front of them? That stands for reversed and effectively reverses the operands. For example

```python
df1.div(df2)
```
would be the same as
```python
df1/df2
```

but....
```python
df1.rdiv(df2)
```
would be the same as
```python
df2/df1
```

### Exercise 3
Create a (3,3) DataFrame and square all elements in it.

In [47]:
pd.DataFrame(np.arange(9).reshape(3,3)) ** 2

Unnamed: 0,0,1,2
0,0,1,4
1,9,16,25
2,36,49,64


### Broadcasting
Similar to numpy, in pandas you can also broadcast data structures. Let's consider a simple example:

In [48]:
df1 = pd.DataFrame(np.arange(16).reshape((4,4)))
df1

Unnamed: 0,0,1,2,3
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11
3,12,13,14,15


In [49]:
df2 = pd.Series(np.arange(4))
df2

0    0
1    1
2    2
3    3
dtype: int64

In [50]:
df1 - df2

Unnamed: 0,0,1,2,3
0,0,0,0,0
1,4,4,4,4
2,8,8,8,8
3,12,12,12,12


Notice how the Series of [0, 1, 2, 3] got removed from each row? That is called broadcasting.

It can also be used for columns, but for that, you have to use the method arithmetic operations.

In [51]:
df1.sub(df2, axis="index")

Unnamed: 0,0,1,2,3
0,0,1,2,3
1,3,4,5,6
2,6,7,8,9
3,9,10,11,12


### Sorting
Sorting is an important built-in operation of pandas. Let's have a look at how you can do it:

In [52]:
df1 = pd.DataFrame(np.arange(16).reshape((4,4)), index=["b", "a", "d", "c"])
df1

Unnamed: 0,0,1,2,3
b,0,1,2,3
a,4,5,6,7
d,8,9,10,11
c,12,13,14,15


In [53]:
df2 = df1.sort_index()
df2

Unnamed: 0,0,1,2,3
a,4,5,6,7
b,0,1,2,3
c,12,13,14,15
d,8,9,10,11


Easy as that. Furthermore, you can also sort along the column axis with
```python
df1.sort_index(axis=1)
```

You can also sort by the actual values inside, but you have to give the column by which you want to sort.

In [54]:
df1 = pd.DataFrame([4, 3, 6, 1, 3, 5], columns=["a"])
df1

Unnamed: 0,a
0,4
1,3
2,6
3,1
4,3
5,5


In [55]:
df1.sort_values(by="a")

Unnamed: 0,a
3,1
1,3
4,3
0,4
5,5
2,6


## Summarizing and computing descriptive stats <a name="sums"></a>
`pandas` is equipped with common mathematical and statistical methods. Most of which fall into the category of reductions or summary statistics. These are methods that extract a single value from a list of values. For example, you can extract the mean of a `Series` object like this:

In [56]:
df = pd.DataFrame(np.arange(20).reshape(5,4),
                 columns=["a", "b", "c", "d"])
df

Unnamed: 0,a,b,c,d
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11
3,12,13,14,15
4,16,17,18,19


In [57]:
df.sum()

a    40
b    45
c    50
d    55
dtype: int64

Notice how that created the sum of each column?

Well you can actually make that the other way around by adding an extra option to `sum()`

In [58]:
df.sum(axis="columns")

0     6
1    22
2    38
3    54
4    70
dtype: int64

A similar method also exists for obtaining the mean of data:

In [59]:
df.mean()

a     8.0
b     9.0
c    10.0
d    11.0
dtype: float64

Finally, the mother of the methods we discussed here is `describe()` 

In [60]:
df.describe()

Unnamed: 0,a,b,c,d
count,5.0,5.0,5.0,5.0
mean,8.0,9.0,10.0,11.0
std,6.324555,6.324555,6.324555,6.324555
min,0.0,1.0,2.0,3.0
25%,4.0,5.0,6.0,7.0
50%,8.0,9.0,10.0,11.0
75%,12.0,13.0,14.0,15.0
max,16.0,17.0,18.0,19.0


Here are some of the summary methods:

| Method | Description |
| -- | -- |
| count          | Return number of non-NA values |
| describe       | Set of summary statistics |
| min, max       | Minimum, maximum values |
| argmin, argmax | Index locations at which the minimum or maximum value is obtained | 
| sum            | Sum of values |
| mean           | Mean of values |
| median         | Arithmetic median of values |
| std            | Sample standard deviation of values
| value_counts() | Counts the number of occurrences of each unique element in a column |

### Exercise 4

A random DataFrame is created below. Find it's mean and standard deviation, then normalise it column-wise according to the formula:

$$ Y = \frac{X - \mu}{\sigma} $$

Where X is your dataset, $\mu$ is the mean and $\sigma$ is the standard deviation.



In [62]:
df = pd.DataFrame(np.random.uniform(0, 10, (100, 100)))

df2 = (df - df.mean())/df.std()
df2

Unnamed: 0,0,1,2,3,4,5,6,7,8,9,...,90,91,92,93,94,95,96,97,98,99
0,0.513870,1.396583,-0.246373,0.412241,1.579367,0.363903,-1.323891,-1.113526,-0.877806,0.881624,...,1.061778,1.647589,0.608297,0.931085,-0.305219,-0.679682,1.323216,-1.000440,1.158411,0.271465
1,1.631868,-1.075335,0.170484,1.453664,1.357602,-1.597306,1.342091,-1.775401,1.129466,-1.180717,...,-1.806966,-0.856221,0.223695,-0.972374,0.176362,0.792399,-0.528541,0.796424,0.728051,-1.104085
2,-0.514866,0.374350,-1.009489,0.086802,-1.696162,0.194635,1.231866,1.121161,0.736265,-0.563802,...,-1.273403,-1.178026,1.542241,1.006310,0.538693,0.789394,0.540420,0.742471,0.203087,-1.486384
3,0.886998,0.052952,1.716351,-0.169000,-1.037321,-1.554467,1.375881,1.256623,-0.206046,-0.299600,...,0.967154,1.519549,1.556544,-0.068761,1.182859,-0.647793,-0.700245,0.705043,-0.825873,1.851311
4,1.391114,0.418370,-1.037966,0.312837,-0.964962,-1.330757,0.641659,1.299796,0.038224,0.356430,...,-0.912055,0.942265,-0.054322,0.174229,-0.782228,-0.613781,-1.241414,0.528770,-1.546610,-0.655011
5,-0.978640,0.882280,1.098095,1.637879,1.623599,1.164122,0.459580,-1.200620,-0.101569,-1.025493,...,-0.178119,0.153881,-1.489814,0.010346,1.043027,1.248067,1.028140,-0.870393,-0.706648,1.744279
6,0.735374,1.603069,-0.607167,1.751502,1.457694,-1.035909,0.547590,0.782969,1.545038,1.538226,...,0.969993,-0.325719,0.994623,-1.531772,1.115576,0.459712,0.949289,-0.787418,-0.295029,0.903614
7,-1.348100,-0.399150,-1.218192,1.062848,0.749450,-1.029018,0.392363,0.421657,0.277607,1.341795,...,0.564476,-1.216358,1.071405,-0.202719,0.165891,-0.592823,0.805909,0.591497,-1.093653,-1.055484
8,1.285660,0.526640,-0.327848,-0.794804,0.136058,0.355371,1.387508,-0.973951,-1.477451,1.033896,...,0.944201,-0.329253,0.670113,0.898907,-1.234780,-1.424890,1.544366,-0.845917,0.405283,0.693724
9,1.564215,-0.530578,-1.378221,1.484833,-0.121855,-1.147937,-0.646704,1.363795,-1.422009,-0.908234,...,-1.839381,0.365275,1.490728,-0.444500,0.555258,1.260674,-0.156228,-1.821950,1.436147,-1.629020


# Data Loading and Storing <a name="loading"></a>
Accessing data is a necessary first step for data science. In this section, the focus will be on data input and output in various formats using `pandas`

Data usually fall into these categories:
- text files
- binary files (more efficient space-wise)
- web data

## Text formats <a name="text"></a>
The most common format in this category is by far `.csv`. This is an easy to read file format which is usually visualised like a spreadsheet. The data itself is usually separated with a `,` which is called the **delimiter**.

Here is an example of a `.csv` file:

```
Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
142, 160, 28, 10, 5, 3,  60, 0.28,  3167
175, 180, 18,  8, 4, 1,  12, 0.43,  4033
129, 132, 13,  6, 3, 1,  41, 0.33,  1471
138, 140, 17,  7, 3, 1,  22, 0.46,  3204
232, 240, 25,  8, 4, 3,   5, 2.05,  3613
135, 140, 18,  7, 4, 3,   9, 0.57,  3028
150, 160, 20,  8, 4, 3,  18, 4.00,  3131
207, 225, 22,  8, 4, 2,  16, 2.22,  5158
271, 285, 30, 10, 5, 2,  30, 0.53,  5702
 89,  90, 10,  5, 3, 1,  43, 0.30,  2054
 ```

It detailed home sale statistics. The first line is called the header, and you can imagine that it is the name of the columns of a spreadsheet.

Let's now see how we can load this data and analyse it. The file is located in the folder `data` and is called `homes.csv`. We can read it like this:

In [63]:
homes = pd.read_csv("data/homes.csv")

In [64]:
homes

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0.28,3167
1,175,180,18,8,4,1,12,0.43,4033
2,143,145,21,7,4,2,10,1.2,3529
3,129,132,13,6,3,1,41,0.33,1471
4,138,140,17,7,3,1,22,0.46,3204
5,232,240,25,8,4,3,5,2.05,3613
6,135,140,18,7,4,3,9,0.57,3028
7,150,160,20,8,4,3,18,4.0,3131
8,293,305,26,8,4,3,6,0.46,7088
9,207,225,22,8,4,2,16,2.22,5158


Easy right?

### Exercise 5
Find the mean selling price of the homes in `data/homes.csv`

In [65]:
homes.mean()

Sell       169.000000
List       176.836364
Living      20.945455
Rooms        7.981818
Beds         3.800000
Baths        1.854545
Age         29.309091
Acres        0.980909
Taxes     3711.545455
dtype: float64

The `read_csv` function has a lot of optional arguments (more than 50). It's impossible to memorise all of them - it's usually best just to look up the particular functionality when you need it. 

You can search `pandas read_csv` online and find all of the documentation.

There are also many other functions that can read textual data. Here are some of them:

| Function | Description
| -- | -- |
| read_csv       | Load delimited data from a file, URL, or file-like object. The default delimiter is a comma `,` |
| read_table     | Load delimited data from a file, URL, or file-like object. The default delimiter is tab `\t` |
| read_clipboard | Reads the last object you have copied (Ctrl-C) |
| read_excel     | Read tabular data from Excel XLS or XLSX file |
| read_hdf       | Read HDF5 file written by pandas |
| read_html      | Read all tables found in the given HTML document |
| read_json      | Read data from a JSON string representation |
| read_sql       | Read the results of a SQL query |

*Note: there are also other loading functions which are not touched upon here*

### Exercise 6
There is another file in the data folder called `homes.xlsx`. Can you read it? Can you spot anything different?

In [66]:
pd.read_excel("data/homes.xlsx")

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142.0,160.0,28.0,10.0,5.0,3.0,60.0,0.28,3167.0
1,175.0,180.0,18.0,8.0,4.0,1.0,12.0,0.43,4033.0
2,129.0,132.0,13.0,6.0,3.0,1.0,41.0,0.33,1471.0
3,138.0,140.0,17.0,7.0,3.0,1.0,22.0,0.46,3204.0
4,232.0,240.0,25.0,8.0,4.0,3.0,5.0,2.05,3613.0
5,135.0,140.0,18.0,7.0,4.0,3.0,9.0,0.57,3028.0
6,150.0,160.0,20.0,8.0,4.0,3.0,18.0,4.0,3131.0
7,207.0,225.0,22.0,8.0,4.0,2.0,16.0,2.22,5158.0
8,271.0,285.0,30.0,10.0,5.0,2.0,30.0,0.53,5702.0
9,89.0,90.0,10.0,5.0,3.0,1.0,43.0,0.3,2054.0


### Writing CSV files
Easy!

In [67]:
homes.to_csv("test.csv")

### Exercise  7
Create a DataFrame which consists of all numbers 0 to 1000. Reshape it into 50 rows and save it to a `.csv` file. How many columns did you end up with?

In [68]:
df = pd.DataFrame(np.arange(1000).reshape(50, -1))
df.to_csv("ex7.csv")

Unnamed: 0,0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19
0,0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19
1,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39
2,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59
3,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79
4,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99
5,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119
6,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139
7,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159
8,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179
9,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199


### Exercise 8
There is a dataset `data/yob2012.txt` which lists the number of newborns registered in 2018 with their names and sex. Open the dataset in pandas **as a csv**, explore it and derive the ratio between male and female newborns.

*Note: The file doesn't contain a header so you will need to add your own column names with*
```python
pd.read_csv("...", names=["Some", "Fun", "Columns"]
```


In [75]:
df = pd.read_csv("data/yob2012.txt", names=["Name", "Sex", "idk"])
sexes = df.Sex.value_counts()

sexes[0] / sexes[1]

1.3694428812605515

## Web scraping <a name="web"></a>
It is also very easy to scrape webpages and extract tables from them.

For example, let's consider extracting the table of failed American banks.

In [76]:
url = "https://www.fdic.gov/bank/individual/failed/banklist.html"
banks = pd.read_html(url)
banks = banks[0]

In [77]:
banks

Unnamed: 0,Bank Name,City,ST,CERT,Acquiring Institution,Closing Date,Updated Date
0,Washington Federal Bank for Savings,Chicago,IL,30570,Royal Savings Bank,"December 15, 2017","February 21, 2018"
1,The Farmers and Merchants State Bank of Argonia,Argonia,KS,17719,Conway Bank,"October 13, 2017","February 21, 2018"
2,Fayette County Bank,Saint Elmo,IL,1802,"United Fidelity Bank, fsb","May 26, 2017","July 26, 2017"
3,"Guaranty Bank, (d/b/a BestBank in Georgia & Mi...",Milwaukee,WI,30003,First-Citizens Bank & Trust Company,"May 5, 2017","March 22, 2018"
4,First NBC Bank,New Orleans,LA,58302,Whitney Bank,"April 28, 2017","December 5, 2017"
5,Proficio Bank,Cottonwood Heights,UT,35495,Cache Valley Bank,"March 3, 2017","March 7, 2018"
6,Seaway Bank and Trust Company,Chicago,IL,19328,State Bank of Texas,"January 27, 2017","May 18, 2017"
7,Harvest Community Bank,Pennsville,NJ,34951,First-Citizens Bank & Trust Company,"January 13, 2017","May 18, 2017"
8,Allied Bank,Mulberry,AR,91,Today's Bank,"September 23, 2016","September 25, 2017"
9,The Woodbury Banking Company,Woodbury,GA,11297,United Bank,"August 19, 2016","December 13, 2018"


Powerful no? Now let's turn that into an exercise.

### Exercise 9
Given the data you just extracted above, can you analyse how many banks have failed per state?

Georgia (GA) should be the state with the most failed banks!

*Hint: try searching the web for pandas counting occurrences* 

In [81]:
banks.ST.value_counts().head()

GA    93
FL    75
IL    69
CA    41
MN    23
Name: ST, dtype: int64

# Data Cleaning <a name="cleaning"></a>
While doing data analysis and modeling, a significant amount of time is spent on data preparation: loading, cleaning, transforming and rearranging. Such tasks are often reported to take **up to 80%** or more of a data analyst's time. Often the way the data is stored in files isn't in the correct format and needs to be modified. Researchers usually do this on an ad-hoc basis using programming languages like Python.

In this chapter, we will discuss tools for handling missing data, duplicate data, string manipulation, and some other analytical data transformations.

## Handling missing data <a name="missing"></a>
Mussing data occurs commonly in many data analysis applications. One of the goals of pandas is to make working with missing data as painless as possible.

In pandas, missing numeric data is represented by `NaN` (Not a Number) and can easily be handled:

In [82]:
string_data = pd.Series(['orange', 'tomato', np.nan, 'avocado'])
string_data

0     orange
1     tomato
2        NaN
3    avocado
dtype: object

In [83]:
string_data.isnull()

0    False
1    False
2     True
3    False
dtype: bool

Furthermore, the pandas `NaN` is functionally equlevant to the standard Python type `NoneType` which can be defined with `x = None`.

In [84]:
string_data[0] = None
string_data

0       None
1     tomato
2        NaN
3    avocado
dtype: object

In [85]:
string_data.isnull()

0     True
1    False
2     True
3    False
dtype: bool

Here are some other methods which you can find useful:
    
| Method | Description |
| -- | -- |
| dropna | Filter axis labels based on whether the values of each label have missing data|
| fillna | Fill in missing data with some value |
| isnull | Return boolean values indicating which values are missing |
| notnull | Negation of isnull |

### Exercise 10
Remove the missing data below using the appropriate method

In [86]:
data = pd.Series([1, None, 3, 4, None, 6])
data.dropna()

0    1.0
2    3.0
3    4.0
5    6.0
dtype: float64

`dropna()` by default removes any row/column that has a missing value. What if we want to remove only rows in which all of the data is missing though?

In [87]:
data = pd.DataFrame([[1., 6.5, 3.], [1., None, None],
                    [None, None, None], [None, 6.5, 3.]])
data

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,1.0,,
2,,,
3,,6.5,3.0


In [88]:
data.dropna()

Unnamed: 0,0,1,2
0,1.0,6.5,3.0


In [89]:
data.dropna(how="all")

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,1.0,,
3,,6.5,3.0


### Exercise 11
That's fine if we want to remove missing data, what if we want to fill in missing data? Do you know of a way? Try to fill in all of the missing values from the data below with **0s**

In [90]:
data = pd.DataFrame([[1., 6.5, 3.], [2., None, None],
                    [None, None, None], [None, 1.5, 9.]])

data.fillna()










pandas also allows us to interpolate the data instead of just filling it with a constant. The easiest way to do that is shown below, but there are more complex ones that are not covered in this course.

In [None]:
data.fillna(method="ffill")

If you want you can explore the other capabilities of [`fillna`](https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.fillna.html)

## Data Transformation  <a name="transformation"></a>
### Removing duplicates
Duplicate data can be a serious issue, luckily pandas offers a simple way to remove duplicates

In [None]:
data = pd.DataFrame([1, 2, 3, 4, 3, 2, 1])
data

In [None]:
data.drop_duplicates()

You can also select which rows to keep

In [None]:
data.drop_duplicates(keep="last")

### Replacing data
You've already seen how you can fill in missing data with `fillna()`. That is actually a special case of more general value replacement. That is done via the `replace()` method.

Let's consider an example where the dataset given to us had `-999` as sentinel values for missing data instead of `NaN`.

In [None]:
data = pd.DataFrame([1., -999., 2., -999., 3., 4., -999, -999, 7.])
data

In [None]:
data.replace(-999, np.nan)

### Detection and Filtering Outliers
Filtering or transforming outliers is largely a matter of applying array operations. Consider a DataFrame with some normally distributed data:

In [4]:
data = pd.DataFrame(np.random.randn(1000, 4))
data.describe()

Unnamed: 0,0,1,2,3
count,1000.0,1000.0,1000.0,1000.0
mean,-0.007632,0.050024,-0.012849,0.016784
std,0.998612,1.006627,0.964622,1.000706
min,-2.957364,-3.502528,-3.227869,-3.169486
25%,-0.665692,-0.626164,-0.662087,-0.67029
50%,0.006478,0.048645,-0.030726,0.04276
75%,0.669843,0.707385,0.645383,0.703211
max,2.772749,3.179202,3.247557,2.902432


Suppose you now want to lower all absolute values exceeding 3 from one of the columns

In [6]:
col = data[2]
col[np.abs(col) > 3]

86     3.247557
211    3.133593
494   -3.227869
Name: 2, dtype: float64

In [7]:
data[np.abs(data) > 3] = np.sign(data) * 3
data.describe()

Unnamed: 0,0,1,2,3
count,1000.0,1000.0,1000.0,1000.0
mean,-0.007632,0.050358,-0.013003,0.016953
std,0.998612,1.004265,0.962654,1.00018
min,-2.957364,-3.0,-3.0,-3.0
25%,-0.665692,-0.626164,-0.662087,-0.67029
50%,0.006478,0.048645,-0.030726,0.04276
75%,0.669843,0.707385,0.645383,0.703211
max,2.772749,3.0,3.0,2.902432


### Exercise 12
Let's load again our file with home prices and filter out homes based on our preference:
1. Load up the file `data/homes.csv`
2. The data contains some duplicates. Filter them out.
3. Let's say that the most we can spend on a house is Â£150. Keep only houses that have a **sell**ing price less than Â£150 and remove the rest
4. Select only houses that have 4 or more bedrooms
5. Select only houses that have 3 or more baths

You should end up with only 2 houses

In [54]:
data = pd.read_csv("data/homes.csv")
data = data.drop_duplicates()
data = data[data["Sell"] < 150]
# data[data["Age"] > 2]
data = data[data["Beds"] >= 4]
data[data["Baths"] >= 3]

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0.28,3167
6,135,140,18,7,4,3,9,0.57,3028
