diff --git a/30_introduction_data_exploration.ipynb b/30_introduction_data_exploration.ipynb index 01ec0236..32fbe7bf 100644 --- a/30_introduction_data_exploration.ipynb +++ b/30_introduction_data_exploration.ipynb @@ -20,6 +20,8 @@ " - [Exercise reading in data](#Exercise-reading-in-data)\n", " - [Playground](#Playground)\n", " - [Data exploration](#Data-exploration)\n", + " - [Guided Exploration Questions](#Guided-Exploration-Questions)\n", + " - [Exercise: Explore the Dataset](#Exercise:-Explore-the-Dataset)\n", " - [Building the plot from scratch](#Building-the-plot-from-scratch)\n", " - [Finding the limits](#Finding-the-limits)\n", " - [Cleaning missing values](#Cleaning-missing-values)\n", @@ -81,6 +83,7 @@ "Python libraries are collections of pre-written code that provide reusable functions and tools for specific tasks.\n", "They significantly extend the capabilities of Python, allowing you to perform complex operations without writing everything from scratch.\n", "This saves time and effort, promotes code reusability, and helps you focus on the higher-level logic of your data analysis and visualization workflows.\n", + "\n", "Have a look at the following plot: imagine you'd have to code every detail from scratch! Instead, we rely on the help of libraries to do most of the heavy lifting for us.\n", "However, there is still quite a bit to do before we can reproduce exactly this.\n", "Let's have a look at the following plot generated with Plotly." @@ -112,7 +115,7 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "In the graph above, the size corresponds to the population of each country and the values of GDP per capital and life expectency along with the name of the country can be seen by hovering over the cursor on the bubbles.\n", + "In the graph above, the size corresponds to each country's population. The values of GDP per capita, life expectancy, and the country's name can be seen by hovering the cursor over the bubbles.\n", "Imagine the work you'd have to put in to build this without any libraries.\n", "\n", "This animated bubble chart can convey a great deal of information since it can accommodate up to *six variables* in total, namely:\n", @@ -126,7 +129,7 @@ "\n", "Using the function `bubbleplot` from the module [`bubbly`(bubble charts with plotly)](https://github.com/AashitaK/bubbly): see references for all source material.\n", "Our goal is to recreate this Visualization but with a different dataset.\n", - "For this, we have already preloaded a data file in the folder `data/introtolibraries/World-happiness-report-updated_2024.csv` which is an open data record which can be found on kaggle.com." + "For this, we have already preloaded a data file in the folder `data/data_exploration/World-happiness-report-updated_2024.csv`, which is an open data record available on kaggle.com." ] }, { @@ -250,13 +253,16 @@ "#### Getting help and inspiration\n", "\n", "Of course one of the most important parts is to be able to understand, look up and get help on any function of a library.\n", - "Usually, we start with some inspiration, as we gave above with the plot, there might be someone who posted something which you would like to reproduce but with a twist or you would like to change something.\n", + "Typically, the creative process begins with an inspiration, such as a plot or an idea you've encountered elsewhere.\n", + "You might wish to reproduce this concept with a unique twist or modify specific elements to suit your vision.\n", "This is generally a good starting point.\n", - "However afterwards you won't have the documentation of all the functions so you need to have the skill to find documentation and understand the requirements for function, sometimes you even need to know more about the inner workings of a functions implementations.\n", "\n", - "There are several ways to access documentation.\n", + "After that, though, you won't have a full documentation guide for every function.\n", + "You'll need to know how to track down the info you need, figure out what each function requires, and sometimes even dig into how they work under the hood.\n", + "\n", + "There are several ways to access a library's documentation.\n", "One way, assuming it is a well-maintained package online, is to find the documentation website.\n", - "For pandas, this is a great place to find details on functions, and changes that may have been made with different versions and explore alternatives to a given function.\n", + "For pandas, this is a great place to find details on functions or changes across different library versions, as well as alternatives to a given function.\n", "\n", "[Pandas Documentation](https://pandas.pydata.org/docs/index.html)\n", "\n", @@ -270,9 +276,9 @@ "And if you are struggling with a specific function you can print out the signature and the docstring with `help(function)` or `function?`.\n", "\n", "When reading the documentation you might get overwhelmed.\n", - "Keep a lookout for the function parameters, many of which may be optional, and good documentations tend to have an example to get a feel for the function.\n", + "Keep a lookout for the function parameters, many of which may be optional.\n", "\n", - "For most of this tutorial we will give you the infos about a function that you need, however if you want to know more or need some extra information then use these tools to inform yourself." + "For most of this tutorial, we will provide you with the information you need about a function. If you want to know more or need some extra information, then use these tools to inform yourself." ] }, { @@ -281,7 +287,7 @@ "source": [ "## First step: Data import and exploration\n", "\n", - "Already getting your data from a file to a variable you can work with can be a headache.\n", + "Getting your data from a file to a variable you can work with can be a headache.\n", "How do I read the file, how do I choose delimiters and what encoding does the file have?\n", "\n", "We will use the `pd.read_csv` function from pandas to read \"The World Happiness Report\" which is a report study on how people rate their happiness in different countries." @@ -295,7 +301,7 @@ "\n", "In the cell below you should write the code that solves the first exercise:\n", "\n", - " - Use the `path_to_happiness` which will be `data/plotly_intro/World-happiness-report-updated_2024.csv` which leads to a CSV file to read in\n", + " - Use the `path_to_happiness`, which will be `data/data_exploration/World-happiness-report-updated_2024.csv`, which leads to a CSV file to read in\n", " - Read in the CSV into a dataframe and output it as `pd.DataFrame`\n", " - Because of how the `.csv`file is formated you must ensure that the encoding is latin1 `encoding='latin1'`" ] @@ -352,8 +358,9 @@ "metadata": {}, "outputs": [], "source": [ - "happyness = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", - "happyness.describe()" + "happiness = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", + "happiness.describe()\n", + "happiness.groupby('year')['Generosity'].mean()" ] }, { @@ -381,7 +388,8 @@ "| `pd.DataFrame['column_name'].value_counts()` | Returns a Series containing counts of unique values in a column. | `df['Status'].value_counts()` |\n", "| `pd.DataFrame.sort_values(by='column_name')` | Sorts the DataFrame by the values in a specified column. | `df.sort_values(by='Date')` |\n", "| `pd.DataFrame.sort_index()` | Sorts the DataFrame by its index. | `df.sort_index()` |\n", - "| `pd.DataFrame.isna().sum()` | Returns the number of missing (NaN) values in each column. | `df.isna().sum()` |\n", + "| `pd.DataFrame.isnull().sum()` | Returns the number of missing (NaN, None) values in each column. | `df.isnull().sum()` |\n", + "| `pd.DataFrame.isna().sum()` | Same as `isnull`, but `isna` is the newer preferred alias. | `df.isna().sum()` |\n", "| `pd.DataFrame.duplicated().sum()` | Returns the number of duplicate rows in the DataFrame. | `df.duplicated().sum()` |\n", "| `pd.DataFrame['column_name'].unique()` | Returns a NumPy array of the unique values in a column. | `df['Country'].unique()` |\n", "| `pd.DataFrame.sample(n=5)` | Returns a random sample of items from the DataFrame (default is 1). | `df.sample(n=10)` |\n", @@ -414,10 +422,8 @@ "source": [ "import pandas as pd\n", "\n", - "happyness = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", + "df = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", "\n", - "# Assuming your dataframe is loaded into 'df'\n", - "df = happyness\n", "print(\"--- First few rows of the dataframe ---\")\n", "print(df.head(2))\n", "print(\"\\n\")\n", @@ -448,8 +454,84 @@ " print(f\"Minimum year: {min_year}\")\n", " print(f\"Maximum year: {max_year}\")\n", " print(f\"Year range: {max_year - min_year} years\")\n", - " print(\"\\n\")\n", - "\n" + " print(\"\\n\")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Guided Exploration Questions\n", + "\n", + "Now that you've seen some basic exploration functions, try answering the following questions about the happiness dataset using the playground cell above.\n", + "These are common questions you would ask when encountering a new dataset for the first time.\n", + "\n", + "1. **How many rows and columns** does the dataset have? (Hint: use `.shape`)\n", + "2. **How many unique countries** are in the dataset? (Hint: use `.nunique()` on the right column)\n", + "3. **Which column has the most missing values?** How many are missing? (Hint: use `.isna().sum()`)\n", + "4. **Are there any duplicate rows?** (Hint: use `.duplicated().sum()`)\n", + "5. **How many data points** does each year have? Is coverage consistent across years? (Hint: use `.groupby('year').size()`)\n", + "6. **Which 5 countries** had the highest `Life Ladder` (happiness) score in 2023? (Hint: filter by year, then use `.nlargest()`)\n", + "7. **What is the mean `Generosity`** across all years? Does it vary much per year? (Hint: use `.groupby('year')['Generosity'].mean()`)\n", + "\n", + "Feel free to experiment with other columns and functions from the reference table above. A good data exploration habit is to always check shape, missing values, duplicates, and distributions before diving into analysis." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Exercise: Explore the Dataset\n", + "\n", + "Now let's formalize some of these exploration steps into a function.\n", + "Complete the function below to return a dictionary with the following keys:\n", + "\n", + "1. `'n_rows'` — the number of rows in the DataFrame (as an `int`).\n", + "2. `'n_columns'` — the number of columns in the DataFrame (as an `int`).\n", + "3. `'n_countries'` — the number of unique countries (as an `int`). Use the `'Country name'` column.\n", + "4. `'n_duplicates'` — the number of duplicate rows (as an `int`).\n", + "5. `'column_with_most_nans'` — the name of the column with the most missing values (as a `str`).\n", + "6. `'happiest_country_2023'` — the name of the country with the highest `'Life Ladder'` score in year 2023 (as a `str`)." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "%reload_ext tutorial.tests.testsuite" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "%%ipytest\n", + "\n", + "import pandas as pd\n", + "import numpy as np\n", + "def solution_explore_dataset(happiness_df: pd.DataFrame) -> dict:\n", + " \"\"\"Explores the happiness dataset and returns key summary statistics.\n", + "\n", + " 1. Count the number of rows and columns using .shape\n", + " 2. Count the number of unique countries in the 'Country name' column\n", + " 3. Count the number of duplicate rows\n", + " 4. Find the column with the most NaN values\n", + " 5. Find the happiest country in 2023 (highest 'Life Ladder')\n", + "\n", + " Args:\n", + " happiness_df : DataFrame containing the happiness data\n", + "\n", + " Returns:\n", + " dict with keys: 'n_rows', 'n_columns', 'n_countries', 'n_duplicates',\n", + " 'column_with_most_nans', 'happiest_country_2023'\n", + " \"\"\"\n", + " # Your code starts here\n", + " return\n", + " # Your code ends here" ] }, { @@ -483,14 +565,19 @@ "## Finding the limits\n", "\n", "If we are plotting a function, it is important to know the order of magnitude of some of the data.\n", - "In our case for example we want to have an animated plot over some years and it helps to know for which years we actually have data.\n", - "In a dataframe we can e.g. use the `.min()` and `.max()` methods. \n", + "For example, in our case, we want to have an animated plot over some years, and it helps to know which years we actually have data.\n", + "To do that in a dataframe we can use the `.min()` and `.max()` methods. \n", "Optionally, to understand the distribution or \"order of magnitude\" of your time values, you might want to plot out the years and check the rough distribution to identify any anomalies or gaps in the data.\n", "This can be done using a histogram or a line plot to visualize the frequency or trend of the time values over the range.\n", "\n", "We want to use `matplotlib.pyplot` for displaying the histogram because it has a useful function hist which does exactly that. \n", + "To do that, we use the `matplotlib.pyplot as plt` library, and there is a `.hist` function which will produce a histogram.\n", "\n", - "We use the `matplotlib.pyplot as plt` library and there there is `.hist` function which will produce a histogram." + "> **Hint:** If the x-axis displays year values as floats (e.g. 2010.0), you can force integer ticks with:\n", + "> ```python\n", + "> from matplotlib.ticker import MaxNLocator\n", + "> plt.gca().xaxis.set_major_locator(MaxNLocator(integer=True))\n", + "> ```" ] }, { @@ -502,12 +589,11 @@ "import pandas as pd\n", "import matplotlib.pyplot as plt\n", "\n", - "happiness = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", - "years = happiness['year'].unique()\n", + "df = pd.read_csv('data/data_exploration/World-happiness-report-updated_2024.csv', encoding='latin1')\n", + "years = df['year'].unique()\n", "print(f\"Unique years in the dataset: {sorted(years)}\")\n", "\n", - "df = happiness\n", - "df['year'] = happiness['year'].astype(int) # Ensure the years are integers\n", + "df['year'] = df['year'].astype(int) # Ensure the years are integers\n", "\n", "# Determine the minimum and maximum years\n", "min_year = df['year'].min()\n", @@ -521,6 +607,11 @@ "plt.xlabel('Year')\n", "plt.ylabel('Frequency')\n", "plt.grid(True)\n", + "\n", + "# Hint: If the x-axis shows float values for years, you can force integer ticks:\n", + "# from matplotlib.ticker import MaxNLocator\n", + "# plt.gca().xaxis.set_major_locator(MaxNLocator(integer=True))\n", + "\n", "plt.show()" ] }, @@ -532,7 +623,7 @@ "\n", "Pandas provides several flexible methods for handling missing data, represented as `NaN`. \n", "You can identify missing values using `.isna()` or `.isnull()`, and then choose a strategy: `.dropna()` removes rows or columns with missing values, while `.fillna()` replaces them. `.ffill` propagates the last valid observation forward to fill in the missing values.\n", - "For example, `df.ffill` will replace a `NaN` with the value from the previous row which had a non-`Nan`value.\n", + "For example, `df.ffill` will replace a `NaN` with the value from the previous row which had a non-`NaN` value.\n", "You can also fill it with a specific value (like the mean, median, or constant).\n", "For time series data, you might use interpolation with `.interpolate()` to fill gaps.\n", "The best approach depends on the nature of the data and the goal of your analysis.\n", @@ -540,7 +631,38 @@ "For this step, we will try to forwardfill the dataframe:\n", "```python\n", " cleaned_happiness = cleaned_happiness.sort_values(by=['Country name', 'year']).ffill()\n", - "```" + "```\n", + "\n", + "> **Hint:** Pandas also provides `.bfill()` (backward fill), which propagates the *next* valid observation backward.\n", + "> While `.ffill()` fills a `NaN` with the value from the previous row, `.bfill()` fills it with the value from the next row.\n", + "> Depending on your data and use case, one may be more appropriate than the other — or you might even combine both to fill gaps from both directions." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "Data cleaning is crucial in data analysis, but several pitfalls exist.\n", + "Here's a summary of common mistakes and how to avoid them:\n", + "\n", + "1. Incorrectly Handling Missing Values.\n", + " Replacing `NaN` values with the mean can be misleading, especially with skewed data.\n", + " Consider using the median or more advanced imputation techniques, and understand the reason for missingness.\n", + " Pandas tools like `fillna()`, `dropna()`, and `interpolate()` are essential here.\n", + "2. Removing Outliers Without Investigation.\n", + " Avoid automatically deleting outliers.\n", + " Visualize the data to determine if outliers are genuine extreme values or errors.\n", + " If genuine, they may be important for analysis. Use boolean indexing with summary statistics to handle them in Pandas.\n", + "3. Ignoring Data Types.\n", + " Ensure columns have the correct data type.\n", + " Use `df.info()` to check and convert columns with `pd.to_numeric()`, `pd.to_datetime()`, or `astype()`.\n", + "4. Not Handling Duplicates Carefully.\n", + " Investigate the source of duplicate rows before removing them.\n", + " They may indicate data entry errors or represent significant repeated measurements.\n", + " Pandas provides `duplicated()` and `drop_duplicates()` for this purpose.\n", + "5. Applying Transformations Incorrectly.\n", + " Scaling data without considering outliers can lead to issues.\n", + " If scaling is necessary, consider robust scalers (like `RobustScaler` from `scikit-learn`) that are less affected by outliers." ] }, { @@ -549,12 +671,12 @@ "source": [ "### Exercise: Complete Happiness\n", "\n", - "In this exercise, we want to complete the dataframe with missing values.\n", - "Complete the function below to \n", + "In this exercise, we want to complete the dataframe's missing values.\n", + "Complete the function below to:\n", "\n", "1. Fill in missing years for every country (so we have an entry for every year between 2005 and 2023 and every country).\n", - " Do this by initializing a DataFrame with `pd.DataFrame()` with a list.\n", - " Then left-merge the happiness dataframe to it with `pd.merge()`.\n", + " First, use a list comprehension to create a list of `(country, year)` tuples for all country/year combinations. Then pass that list to `pd.DataFrame()` with `columns=['Country name', 'year']` to build a complete scaffold DataFrame.\n", + " Finally, use the `.merge()` method on this scaffold to left-merge the original happiness data onto it (matching on `'Country name'` and `'year'`). This ensures every country has a row for every year, with `NaN` where data was missing.\n", "2. Fill all missing values in the year 2005 with the value 1.\n", " Use the `.fillna()` function.\n", "3. Forwardfill all the remaining years with the function `.ffill()`.\n", @@ -583,7 +705,7 @@ "def solution_clean_dataset(happiness_df: pd.DataFrame) -> pd.DataFrame:\n", " \"\"\"Cleans the dataset by adding missing year and country values\n", "\n", - " 1. Add in missing years for every country\n", + " 1. Create a DataFrame with all (country, year) combinations, then left-merge the happiness data onto it\n", " 2. Fill the minimum year with values of 1\n", " 3. Forward fill the rest of the years\n", "\n", @@ -598,33 +720,6 @@ " # Your code ends here" ] }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "Data cleaning is crucial in data analysis, but several pitfalls exist.\n", - "Here's a summary of common mistakes and how to avoid them:\n", - "\n", - "1. Incorrectly Handling Missing Values.\n", - " Replacing `NaN` values with the mean can be misleading, especially with skewed data.\n", - " Consider using the median or more advanced imputation techniques, and understand the reason for missingness.\n", - " Pandas tools like `fillna()`, `dropna()`, and `interpolate()` are essential here.\n", - "2. Removing Outliers Without Investigation.\n", - " Avoid automatically deleting outliers.\n", - " Visualize the data to determine if outliers are genuine extreme values or errors.\n", - " If genuine, they may be important for analysis. Use boolean indexing with summary statistics to handle them in Pandas.\n", - "3. Ignoring Data Types.\n", - " Ensure columns have the correct data type.\n", - " Use `df.info()` to check and convert columns with `pd.to_numeric()`, `pd.to_datetime()`, or `astype()`.\n", - "4. Not Handling Duplicates Carefully.\n", - " Investigate the source of duplicate rows before removing them.\n", - " They may indicate data entry errors or represent significant repeated measurements.\n", - " Pandas provides `duplicated()` and `drop_duplicates()` for this purpose.\n", - "5. Applying Transformations Incorrectly.\n", - " Scaling data without considering outliers can lead to issues.\n", - " If scaling is necessary, consider robust scalers (like `RobustScaler` from `scikit-learn`) that are less affected by outliers." - ] - }, { "cell_type": "markdown", "metadata": {}, @@ -635,8 +730,8 @@ "To add a regional indicator, we'll need another dataset that maps countries to their respective regions.\n", "We can then merge this data with our main happiness dataframe based on the 'Country name' column.\n", "\n", - "Let's assume you have a CSV file named `country_region_mapping.csv` in your `data/plotly_intro/` directory with columns `'Country name'` and `'Region indicator'`.\n", - "Then we could merge the dataframes on the country name and get a regional indicator for all the countries.\n", + "We have a CSV file named `country_region_mapping.csv` in the `data/data_exploration/` directory with columns `'Country name'` and `'Regional indicator'`.\n", + "We can merge the dataframes on the country name and get a regional indicator for all the countries.\n", "\n", "Let's explore the `pd.merge` function for that.\n", "The merging of tables comes from SQL Table merges and if you are not familiar with those, for now, keep in mind we want to do a **left merge** with the happiness table being the left table and the region mapping the right table and we merge **on** a column which they have in common.\n", @@ -697,8 +792,8 @@ "def solution_add_regional_indicator(cleaned_happiness_df: pd.DataFrame, region_df: pd.DataFrame) -> pd.DataFrame:\n", " \"\"\"Adds a regional indicator to the dataset\n", "\n", - " 1. Merge the cleaned_happiness_df with region_df on the 'Country name' and 'year' columns\n", - " 2. Fill the missing values in the 'Region indicator' column with 'Unknown'\n", + " 1. Merge the cleaned_happiness_df with region_df on the 'Country name' column\n", + " 2. Fill the missing values in the 'Regional indicator' column with 'Unknown'\n", "\n", " Args:\n", " cleaned_happiness_df : DataFrame containing the happiness data\n", @@ -743,7 +838,8 @@ "\n", "The data can be seen as the plot data initially, and the frames are then the animation steps.\n", "\n", - "Let's first try to create a simple scatter plot, for that we populate a figure dictionary data with a trace (`dict`) which contains an array of values for `x`, an array of values for `y`, a `mode` ('markers') and an array of strings for the text which is what appears when hovered over.\n", + "Let's first try to create a simple scatter plot.\n", + "For that, we populate a figure dictionary data with a trace (`dict`) which contains an array of values for `x`, an array of values for `y`, a `mode` ('markers') and an array of strings for the text, which is what appears when hovered over.\n", "\n", "```python\n", "trace = {\n", @@ -771,7 +867,6 @@ "x_column = 'Freedom to make life choices'\n", "y_column = 'Life Ladder'\n", "description_column = 'Country name'\n", - "# time_column = 'year'\n", "\n", "\n", "\n", @@ -829,7 +924,6 @@ "x_column = 'Freedom to make life choices'\n", "y_column = 'Life Ladder'\n", "description_column = 'Country name'\n", - "# time_column = 'year'\n", "figure = get_scatter_figure(dataset, x_column, y_column, description_column)\n", "\n", "def frame_by_year(dataset, year, x_column, y_column, description_column):\n", @@ -867,7 +961,8 @@ "source": [ "### Adding slider bar for time scale\n", "\n", - "The slider needs configuring, this would require a bit of reading up what exactly you need or if you have an example you can make use of the existing functions.\n", + "The slider needs configuring.\n", + "This would require a bit of reading up on what exactly you need, or, if you have an example, you can make use of the existing functions.\n", "The following is heavily inspired by the module [`bubbly`](https://github.com/AashitaK/bubbly).\n", "This is simply a configuration and contains only the years data." ] @@ -1047,7 +1142,7 @@ "# append to final_happiness_df\n", "dataset = pd.merge(complete_happiness_df, resized_log_gdp_df, on=['Country name', 'year'], how='left')\n", "# Check out year 2010\n", - "dataset[dataset['year'] == 2010].head(10)\n" + "dataset[dataset['year'] == 2010].head(10)" ] }, { @@ -1069,22 +1164,20 @@ "y_column = 'Life Ladder'\n", "description_column = 'Country name'\n", "time_column = 'year'\n", - "# Set the layout\n", - "figure = set_layout(x_title='Freedom to make life choices', y_title='Life Ladder',\n", - " title='Happiness Indicators', x_logscale=False, y_logscale=False, \n", - " show_slider=True, slider_scale=years, show_button=True, show_legend=False, \n", - " height=650)\n", "\n", "# Define the new variable\n", "bubble_size_column = 'Resized Log GDP per capita'\n", "category_column = 'Regional indicator'\n", "\n", - "\n", - "\n", "# Make the grid\n", "years = dataset[time_column].unique()\n", "years.sort()\n", - " \n", + "\n", + "# Set the layout\n", + "figure = set_layout(x_title='Freedom to make life choices', y_title='Life Ladder',\n", + " title='Happiness Indicators', x_logscale=False, y_logscale=False, \n", + " show_slider=True, slider_scale=years, show_button=True, show_legend=False, \n", + " height=650)\n", "\n", "# Add the base frame\n", "year = min(years)\n", @@ -1208,10 +1301,12 @@ " year : Year to plot\n", " x_column : Column name for x-axis\n", " y_column : Column name for y-axis\n", - " description_column : Column name for text\n", + " description_column : Column name for hover text\n", + " category_column : Column name for the category to split traces by\n", + " bubble_size_column : Column name for the bubble size\n", "\n", " Returns:\n", - " - Dictionary containing the trace and frame information\n", + " - Dictionary with 'data' (list of traces per category) and 'name' (year as string)\n", " \"\"\"\n", " # Your code starts here\n", " return\n", @@ -1252,23 +1347,31 @@ "However, if you are browsing through possible libraries to use you might also find less well-maintained libraries, ones that may only have a single author and ones that haven't been touched in a while. \n", "\n", "Here we give you a direct example, this tutorial was inspired by the [bubbly](https://github.com/AashitaK/bubbly) package.\n", - "However, with an update from the pandas library it is no longer compatible with newer versions of pandas and will through an error (see codeblock below). \n", + "However, with an update from the pandas library, it is no longer compatible with newer versions of pandas and will throw an error (see codeblock below). \n", "So what to do in that case?\n", "\n", - "There are many options, you can inform the author of this problem on GitHub.\n", - "Of course, they may not have time to fix this.\n", + "There are many options.\n", + "For example, you can inform the author of this problem on GitHub.\n", + "But of course, they may not have time to fix this.\n", "You can find a different library, however, it might not be exactly the way you wanted it.\n", - "You can downgrade your pandas library to be compatible, if you use pip show pandas you will see what version you have, it is possible to uninstall and reinstall a specific version.\n", - "However, this might not be feasible if you need it in other places and is generally not a pretty solution. \n", + "You can downgrade your pandas library to be compatible.\n", + "If you use `pip show pandas`, you will see what version you have, and it is possible to uninstall and reinstall a specific version.\n", + "However, this might not be feasible if you need it in other places and is generally not the best solution. \n", "Last but not least you can try to fix it yourself.\n", "\n", - "So as an exercise, we exported the bubbly library as a file `bubbly.py` into the folder `data.plotly_intro`.\n", + "So as an exercise, we exported the bubbly library as a file `bubbly.py` into the folder `data/data_exploration`.\n", "It is quite a short library so quite managable.\n", - "Try to figure out what the error is exactly and then fix the library locally by modifying only the file `data/plotly_intro/bubbly.py` until the same code below compiles.\n", + "Try to figure out what the error is exactly and then fix the library locally by modifying only the file `data/data_exploration/bubbly.py` until the same code below compiles.\n", + "\n", + "Note: You will need to restart the kernel after applying changes to the packages.\n", "\n", - "Note: You will need to restart the kernel after changes to the packages.\n", + "
\n", + "Hint (click to reveal)\n", "\n", - "(If you are interested in a solution, we have a fixed version under tutorial.my_bubbly.py, feel free to check the differences.)" + "The changes are all in the `make_grid` / `make_grid_with_categories` functions around lines 272, 282, 313, and 328.\n", + "The broken code uses a pandas method that was removed in newer versions.\n", + "\n", + "
" ] }, { @@ -1290,6 +1393,23 @@ " x_logscale=True, scale_bubble=3, height=650)\n", "iplot(figure, config={'scrollZoom': True})\n" ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "If you are interested in a solution, we have a fixed version under `tutorial/my_bubbly.py`. You can compare the two files with Python's `difflib` module:\n", + "\n", + "```python\n", + "import difflib\n", + "with open('data/data_exploration/bubbly.py') as f:\n", + " original = f.readlines()\n", + "with open('tutorial/my_bubbly.py') as f:\n", + " fixed = f.readlines()\n", + "diff = difflib.unified_diff(original, fixed, fromfile='bubbly.py', tofile='my_bubbly.py')\n", + "print(''.join(diff))\n", + "```" + ] } ], "metadata": { diff --git a/tutorial/data_exploration_helper.py b/tutorial/data_exploration_helper.py index 04080e68..e2062193 100644 --- a/tutorial/data_exploration_helper.py +++ b/tutorial/data_exploration_helper.py @@ -380,6 +380,7 @@ def set_layout( figure["layout"]["title"] = title figure["layout"]["hovermode"] = "closest" figure["layout"]["showlegend"] = show_legend + figure["layout"]["legend"] = {"itemsizing": "constant"} figure["layout"]["margin"] = {"b": 50, "t": 50, "pad": 5} if width: diff --git a/tutorial/tests/test_30_introduction_data_exploration.py b/tutorial/tests/test_30_introduction_data_exploration.py index b9a635cc..decf5050 100644 --- a/tutorial/tests/test_30_introduction_data_exploration.py +++ b/tutorial/tests/test_30_introduction_data_exploration.py @@ -29,8 +29,49 @@ def test_read_in_dataframe(input_arg, function_to_test): # Read in the data happiness_df = reference_read_in_dataframe(path_to_happiness) + result = function_to_test(path_to_happiness) + assert isinstance(result, pd.DataFrame), ( + "Your function should return a pd.DataFrame, but it returned None. Did you forget the return statement?" + ) # Check if the two DataFrames are equal - assert happiness_df.equals(function_to_test(path_to_happiness)) + assert happiness_df.equals(result) + + +def reference_explore_dataset(happiness_df: pd.DataFrame) -> dict: + n_rows, n_columns = happiness_df.shape + n_countries = happiness_df["Country name"].nunique() + n_duplicates = int(happiness_df.duplicated().sum()) + column_with_most_nans = happiness_df.isna().sum().idxmax() + happiest_2023 = ( + happiness_df[happiness_df["year"] == 2023] + .nlargest(1, "Life Ladder")["Country name"] + .iloc[0] + ) + return { + "n_rows": n_rows, + "n_columns": n_columns, + "n_countries": n_countries, + "n_duplicates": n_duplicates, + "column_with_most_nans": column_with_most_nans, + "happiest_country_2023": happiest_2023, + } + + +@pytest.mark.parametrize("input_arg", input_args) +def test_explore_dataset(input_arg, function_to_test): + happiness_df = reference_read_in_dataframe( + "data/data_exploration/World-happiness-report-updated_2024.csv" + ) + ref = reference_explore_dataset(happiness_df) + sol = function_to_test(happiness_df) + assert isinstance(sol, dict), ( + "Your function should return a dict, but it returned None. Did you forget the return statement?" + ) + for key in ref: + assert key in sol, f"Missing key '{key}' in the returned dictionary" + assert sol[key] == ref[key], ( + f"Value for '{key}' is {sol[key]}, expected {ref[key]}" + ) def reference_clean_dataset(happiness_df: pd.DataFrame) -> pd.DataFrame: @@ -71,6 +112,12 @@ def test_clean_dataset(input_arg, function_to_test): clean_ref = reference_clean_dataset(hapiness_df) clean_sol = function_to_test(hapiness_df) + assert isinstance(clean_sol, pd.DataFrame), ( + "Your function should return a pd.DataFrame, but it returned None. Did you forget the return statement?" + ) + assert "Country name" in clean_sol.columns and "year" in clean_sol.columns, ( + "The output should contain 'Country name' and 'year' columns" + ) clean_ref_sorted = clean_ref.sort_values(by=["Country name", "year"]).reset_index( drop=True ) @@ -126,8 +173,14 @@ def test_add_regional_indicator(input_arg, function_to_test): clean_ref = reference_add_regional_indicator(cleaned_happiness_df, region_df) clean_sol = function_to_test(cleaned_happiness_df, region_df) - # Check if the two DataFrames are equal - assert clean_ref.equals(clean_sol) + assert isinstance(clean_sol, pd.DataFrame), ( + "Your function should return a pd.DataFrame, but it returned None. Did you forget the return statement?" + ) + # Sort both DataFrames to ensure order-independent comparison + sort_cols = ["Country name", "year"] + clean_ref_sorted = clean_ref.sort_values(by=sort_cols).reset_index(drop=True) + clean_sol_sorted = clean_sol.sort_values(by=sort_cols).reset_index(drop=True) + assert clean_ref_sorted.equals(clean_sol_sorted) # solution_frames_with_category @@ -237,8 +290,26 @@ def test_frames_with_category(input_arg, function_to_test): bubble_size_column, ) - # Check if the two DataFrames are equal - assert clean_ref == clean_sol + assert isinstance(clean_sol, dict), ( + "Your function should return a dict, but it returned None. Did you forget the return statement?" + ) + assert "data" in clean_sol and "name" in clean_sol, ( + "The returned dict should have 'data' and 'name' keys" + ) + assert clean_ref["name"] == clean_sol["name"], ( + f"Frame name mismatch: expected '{clean_ref['name']}', got '{clean_sol['name']}'" + ) + # Compare traces regardless of category order + ref_traces = sorted(clean_ref["data"], key=lambda t: t["name"]) + sol_traces = sorted(clean_sol["data"], key=lambda t: t["name"]) + assert len(ref_traces) == len(sol_traces), ( + f"Expected {len(ref_traces)} traces, got {len(sol_traces)}" + ) + for ref_t, sol_t in zip(ref_traces, sol_traces, strict=True): + assert ref_t["name"] == sol_t["name"], ( + f"Trace name mismatch: expected '{ref_t['name']}', got '{sol_t['name']}'" + ) + assert ref_t == sol_t, f"Trace data mismatch for category '{ref_t['name']}'" from plotly.offline import iplot