Skip to content

Commit 30d3ef8

Browse files
committed
Final edits
1 parent 3a1bf70 commit 30d3ef8

1 file changed

Lines changed: 8 additions & 6 deletions

File tree

2024-landscape.md

Lines changed: 8 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -9,12 +9,14 @@ site:
99
Python is widely adopted in data science, and its use for statistics is expanding rapidly---particularly in education and applied research.
1010
The statistical ecosystem in Python is currently anchored by six major libraries:
1111

12-
- [numpy](https://www.numpy.org/), which provides fast, flexible array and numerical operations and underpins nearly all statistical and scientific computing in Python.
13-
- [pandas](https://www.pandas.org/), which offers intuitive, high-performance data structures for tabular and time series data, making data cleaning, wrangling, and exploration straightforward and efficient.
14-
- [scipy.stats](https://docs.scipy.org/doc/scipy/reference/stats.html), which provides a comprehensive suite of probability distributions, summary statistics, and basic statistical tests.
15-
- [statsmodels](https://www.statsmodels.org/), which offers tools for econometrics, classical statistics, and statistical modeling---including linear and generalized linear models, time series analysis, and hypothesis testing.
16-
- [scikit-learn](https://scikit-learn.org/), which is best known for machine learning but also supports some statistical modeling, offering a consistent API for predictive analytics and data preprocessing.
17-
- [seaborn](https://seaborn.pydata.org/), a library built on top of matplotlib that excels at creating informative and attractive statistical graphics, making it easier to visualize distributions, relationships, and trends in data.
12+
- [numpy](https://www.numpy.org/), which provides fast, flexible array and numerical operations, and underpins nearly all statistical and scientific computing in Python.
13+
It supports descriptive statistics, correlation and covariance computations, random sampling, and tools for constructing histograms and binning data.
14+
- [pandas](https://www.pandas.org/), which offers intuitive, high-performance data structures for tabular and time series data, making data cleaning, wrangling, reshaping, aggregation, and exploratory analysis straightforward and efficient.
15+
- [scipy](https://www.scipy.org/), which builds on NumPy to deliver a broad range of scientific and statistical functionality---including, in its [`scipy.stats`](https://docs.scipy.org/doc/scipy/reference/stats.html) submodule, a comprehensive suite of probability distributions, summary statistics, and basic statistical tests.
16+
It also provides modules for clustering, optimization, interpolation, and signal processing.
17+
- [matplotlib](https://matplotlib.org/), the foundational plotting library in Python, which enables the creation of high-quality static, animated, and interactive visualizations, and serves as the basis for many higher-level plotting and statistical graphics libraries.
18+
- [statsmodels](https://www.statsmodels.org/), which offers tools for econometrics, classical statistics, and statistical modeling---including linear and generalized linear models, time series analysis, survival analysis, and hypothesis testing, with extensive support for model diagnostics and statistical inference.
19+
- [scikit-learn](https://scikit-learn.org/), which is best known for machine learning but also supports statistical modeling, offering a consistent API for regression, classification, clustering, model evaluation, statistical preprocessing, and dimensionality reduction.
1820

1921
These core libraries are generally well-tested, reliable, and uphold high software engineering standards, making them trusted foundations for research and application.
2022
They benefit from contributions not only from science users but also from methods and software developers.

0 commit comments

Comments
 (0)