Improve Data Exploration notebook - #368
Conversation
- Updated the path for the happiness CSV file in the data exploration notebook to reflect the correct directory structure. - Added hints for visualizing year values in matplotlib plots. - Improved documentation for data cleaning functions, clarifying the merging process and filling strategies. - Enhanced test cases to assert the return types and structure of DataFrames in the testing suite for better validation. - Adjusted layout settings in the helper module to ensure consistent legend sizing in visualizations.
for more information, see https://pre-commit.ci
There was a problem hiding this comment.
Pull request overview
Updates the “Introduction: Data exploration” tutorial materials by expanding the learning content (new guided questions + a new exercise) and improving robustness/clarity of the associated autograder tests and plotting helper behavior.
Changes:
- Added a new “Explore the Dataset” exercise and corresponding pytest coverage.
- Improved existing tests with clearer return-type assertions and order-independent comparisons.
- Tweaked Plotly layout helper to keep legend item sizing constant.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
tutorial/tests/test_30_introduction_data_exploration.py |
Adds a new test_explore_dataset and strengthens several existing tests (type assertions, stable comparisons). |
tutorial/data_exploration_helper.py |
Adjusts Plotly layout to set legend itemsizing to "constant". |
30_introduction_data_exploration.ipynb |
Adds guided exploration + a new exercise section and fixes/rewrites various tutorial text and paths. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| "You can identify missing values using `.isna()` or `.isnull()`, and then choose a strategy: `.dropna()` removes rows or columns with missing values, while `.fillna()` replaces them. `.ffill` propagates the last valid observation forward to fill in the missing values.\n", | ||
| "For example, `df.ffill` will replace a `NaN` with the value from the previous row which had a non-`Nan`value.\n", | ||
| "For example, `df.ffill` will replace a `NaN` with the value from the previous row which had a non-`NaN` value.\n", |
There was a problem hiding this comment.
This sentence refers to df.ffill (the method object) rather than calling it; as written it’s not accurate. Consider changing it to df.ffill() (and similarly keep .ffill() consistent throughout the paragraph).
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
yakutovicha
left a comment
There was a problem hiding this comment.
A bit of nitpicking, feel free to take or ignore.
Co-authored-by: Aliaksandr Yakutovich <aliaksandr.yakutovich@gmail.com>
fixes #354
Fixing comments from the issue to the data_eploration notebook. Adding an exercise and fixing a bunch of typos.