# Notebook 2 - pandas
[pandas](http://pandas.pydata.org) provides high-level data structures and functions designed to make working with structured or tabular data fast, easy and expressive. The primary objects in pandas that we will be using are the `DataFrame`, a tabular, column-oriented data structure with both row and column labels, and the `Series`, a one-dimensional labeled array object.

pandas blends the high-performance, array-computing ideas of NumPy with the flexible data manipulation capabilities of spreadsheets and relational databases. It provides sophisticated indexing functinoality to make it easy to reshape , slice and perform aggregations.

While pandas adopts many coding idioms from NumPy, the biggest difference is that pandas is designed for working with tabular or heterogeneous data. NumPy, by contrast, is best suited for working with homogeneous numerical array data.

<br>
**Table of Contents:**
- Data Structures
    - Series
    - DataFrame
- Essential Functionality
    - Reindexing
    - Dropping Entries
    - Indexing, Slicing and Filtering
    - Arithmetic Operations
    - Sorting and ranking
- Summarizing and Computing Descriptive Statistics
    - Correlation and Covariance
    - Unique values, value counts and Membership
- Reading and storing data
    - Text Format
    - Text Format Writing
    - XML and HTML Web Scraping
    - Reading excel files
    - mention that pandas allow interfacing with web APIs and SQL databases
- Data Cleaning and preperation
    - Missing data
    - Data transformation
    - String manipulation incl. regexp
- Data wrangling
- Plotting?
    

In [2]:
# Common pandas import statement
import pandas as pd

# Data Structures
## Series
a one-dimensional array-like object containing a sequence of values and an associated array of data labels, called its index.

The easiest way to make a Series is from an array of data:

In [4]:
data = pd.Series([4, 7, -5, 3])

Now try printing out data

The string representation of a Seires displayed interactively shows the index on the left and the values on the right. Because we didn't specify an index, the default on is simply integers 0 through N-1.

You can output only the values of a Series using 
```python
data.values
```
or you can get only the indeces using
```python
data.index
```
Try it out below!

You can specify custom indeces when intialising the Series

In [7]:
data2 = pd.Series([4, 7, -5, 3], index=["a", "b", "c", "d"])

Now you can use these labels to access the data similar to a normal array

In [8]:
data2["a"]

4

Another way to think about Serieses is as a fixed-length ordered dictionary. Furthermore, you can actually define a Series in a similar manner to a dictionary

In [10]:
cities = {"Glasgow" : 599650, "Edinburgh" : 464990, "Abardeen" : 196670, "Dundee" : 147710}
data3 = pd.Series(cities)

In [11]:
data3

Abardeen     196670
Dundee       147710
Edinburgh    464990
Glasgow      599650
dtype: int64

You can do arithmetic operations between Serieses similar to NumPy arrays. Even if you have 2 datasets with different data, arithmetic operations will be aligned according to their indeces.

Let's look at an example

In [12]:
cities_uk = {"Birmingham" : 1092330, "Leeds": 751485, "Glasgow" : 599650,
             "Manchester" : 503127, "Edinburgh" : 464990}
data4 = pd.Series(cities_uk)

In [13]:
data3 + data4

Abardeen            NaN
Birmingham          NaN
Dundee              NaN
Edinburgh      929980.0
Glasgow       1199300.0
Leeds               NaN
Manchester          NaN
dtype: float64

Notice how some of the results are NaN? Well that is because there were no instances of those cities within both of the datasets. You can usually extract NaNs from a Series with
```python
data4.isnull()
```

## DataFrame
A DataFrame represents a rectangular table of data and contains an ordered collection of columns, each of which can be a different value type. The DataFrame has both row and column index and can be thought of as a dict of Series all sharing the same index.

The most common way to create a DataFrame is with dicts

In [17]:
data = {"cities" : ["Glasgow", "Edinburgh", "Abardeen", "Dundee"],
        "population" : [599650, 464990, 196670, 147710],
        "year" : [2011, 2013, 2013, 2013]}
frame = pd.DataFrame(data)

Try printing it out

In [15]:
frame

Unnamed: 0,cities,population,year
0,Glasgow,599650,2011
1,Edinburgh,464990,2013
2,Abardeen,196670,2013
3,Dundee,147710,2013


Jupyter Notebooks prints it out in a nice table but the basic version of this is also just as readable!

Additionally you can also specify the order of columns during initialisation

In [18]:
frame2 = pd.DataFrame(data, columns=["year", "cities", "population"])

You can retrieve a particular column from a DataFrame with
```python
frame["cities"]
```
The result is going to be a Series

Additionally you can retrieve a row from the dataset using
```python
frame[1]
```

It is also possible to add and modify the columns of a DataFrame

In [20]:
frame2["size"] = 100

In [21]:
frame2

Unnamed: 0,year,cities,population,size
0,2011,Glasgow,599650,100
1,2013,Edinburgh,464990,100
2,2013,Abardeen,196670,100
3,2013,Dundee,147710,100


In [22]:
frame2["size"] = [175, 264, 65.1, 60]  # in km^2

Similar to dicts, columns can be deleted using
```python
del frame2["size"]
```

Another common way of creating DataFrames is from a nested dict of dicts:

In [None]:
data2 = {"cities" : ["Glasgow", "Edinburgh", "Abardeen", "Dundee"],
        "population" : [599650, 464990, 196670, 147710],
        "year" : [2011, 2013, 2013, 2013]}

data2 = {"Glasgow": {2011: 599650},
        "Edinburgh": {2013:464990},
        "Abardeen": }

frame3 = pd.DataFrame(data)

Here is a table of different ways of intialising a DataFrame for your reference

| Type | Notes |
| --- | --- |
| 2D ndarray | A matrix of data; passing optional row and column labels |
| dict of arrays, lists, or tuples | Each sequence becomes a column in the DataFrame; all sequences must be the same length |
| NumPy structured/recorded array | Treated as with the "dict of arrays, lists or tuples" case |
| dict of Serires | Each value becomes a column; indexes from each Series are unioned together to<br>form the result's row index if not explicit index is passed |
| dict of dicts | Each inner dict becomes a column; keys are unioned to form the row<br>index as in the "dict of Series" case |
| List of dicts or Series | Each item becomes a row in the DataFrame; union of dict keys or<br>Series indeces become the DataFrame's column labels |
| List of lists or tuples | Treated as the "2D ndarray" case |
| Another DataFrame | The DataFrame's indexes are used unless different ones are passed |
| NumPy MaskedArray | Like the "2D ndarray" case except masked values become NA/missing in the DataFrame |

# Essential Functionality
In this section we will go through the fundemental mechanics of interacting with the data contained in a Series or DaraFrame.

## Reindexing