scikit-learn: The Machine Learning Library That Isn't Deep Learning

scikit-learn is a free, open-source machine learning library for Python — and, according to a 2022 Kaggle survey of nearly 24,000 developers across 173 countries, the single most widely used machine learning framework, ahead of both TensorFlow and PyTorch. It covers classification, regression, and clustering with well-established algorithms like random forests, gradient boosting, support-vector machines, k-means, and DBSCAN, and it’s licensed under the permissive New BSD License. Current stable release is 1.9.1.
Unlike the other three frameworks in this series, scikit-learn isn’t a deep learning library — it doesn’t train neural networks. It’s the tool most people reach for before or alongside one: cleaning data, engineering features, running a baseline model, and evaluating results, all through one deliberately consistent API. Every fact here is checked against scikit-learn’s own documentation, its changelog, and Wikipedia’s sourced entry, and the code is copied from scikit-learn’s own current getting-started guide.
From a summer project to the most-used ML library
scikit-learn’s origin is smaller than the other three frameworks in this series: it started as scikits.learn, a Google Summer of Code project by French data scientist David Cournapeau in June 2007 — originally a third-party extension to SciPy, hence the name (“scientific toolkit for machine learning”). Later that year, Matthieu Brucher joined and began using it in his own thesis work.
The project changed hands in 2010: INRIA, the French Institute for Research in Computer Science and Automation, took over leadership through contributors Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, and Vincent Michel, and published the first public release (v0.1 beta) in late January 2010. It stayed a fast-moving, pre-1.0 project for over a decade — the first stable 1.0.0 release didn’t arrive until 24 September 2021, the result of more than 2,100 merged pull requests, roughly 800 of them documentation work alone.
scikit-learn is a NumFOCUS fiscally sponsored project — the same nonprofit umbrella that backs NumPy, pandas, and Matplotlib — which is a smaller, quieter version of the governance story running through this whole series: PyTorch moved to the Linux Foundation in 2022, TensorFlow’s XLA compiler opened into OpenXLA in 2023, and scikit-learn has been under independent, non-corporate stewardship for most of its life.
The idea that makes it different: one consistent API
scikit-learn’s actual design philosophy, not a marketing claim: every model, called an estimator, is fit and used the same way, regardless of whether it’s a random forest, a linear model, or a clustering algorithm. This is the library’s real competitive advantage over hand-rolling algorithms individually — you learn the pattern once:
from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier(random_state=0)
X = [[1, 2, 3], # 2 samples, 3 features
[11, 12, 13]]
y = [0, 1] # class of each sample
clf.fit(X, y)
clf.predict(X) # array([0, 1])
clf.predict([[4, 5, 6]]) # predicts on new data, no retraining
Preprocessing steps — scaling, encoding, imputing missing values — follow the identical shape, just with .transform() instead of .predict():
from sklearn.preprocessing import StandardScaler
X = [[0, 15], [1, -10]]
StandardScaler().fit(X).transform(X)
# array([[-1., 1.],
# [ 1., -1.]])
A real pipeline, end to end
Because preprocessors and models share the same interface, scikit-learn can chain them into a single Pipeline object that behaves exactly like any other estimator — fit once, predict once, and never accidentally leak test data into training. This is scikit-learn’s own current documentation example, training a classifier on the Iris dataset:
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
pipe = make_pipeline(StandardScaler(), LogisticRegression())
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)
pipe.fit(X_train, y_train)
accuracy_score(pipe.predict(X_test), y_test)
# 0.97...
Model evaluation is built in rather than left to you to script by hand. A 5-fold cross-validation, for instance, is one call:
from sklearn.datasets import make_regression
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_validate
X, y = make_regression(n_samples=1000, random_state=0)
result = cross_validate(LinearRegression(), X, y) # defaults to 5-fold CV
result['test_score']
And hyperparameter tuning — finding the best n_estimators or max_depth for a random forest, for example — is handled by RandomizedSearchCV or GridSearchCV, which wrap any estimator and search its parameter space using the same cross-validation machinery under the hood.
2025’s quiet addition: GPU support, through other frameworks’ arrays
scikit-learn has always run on CPU, backed by NumPy and SciPy — with performance-critical algorithms like support-vector machines and logistic regression implemented in Cython, wrapping the C libraries LIBSVM and LIBLINEAR directly rather than reimplementing them in Python. That changed, carefully and incrementally, in version 1.8 (December 2025): scikit-learn began adding Array API support to individual estimators — StandardScaler, RidgeCV, confusion_matrix, roc_curve, and dozens of others, with more added every release.
The Array API is a standard interface that NumPy, PyTorch, and CuPy all implement. An estimator written against it doesn’t just accept a NumPy array anymore — it accepts a PyTorch tensor or a CuPy array too, and runs its computation on whatever device that array already lives on. Pass in a CUDA-backed PyTorch tensor, and a supported estimator computes on the GPU, with no separate GPU-specific code path to install or call. scikit-learn’s own documentation is explicit that this is still experimental and rolling out estimator by estimator, not a library-wide switch — check a given estimator’s own documentation before assuming it applies.
The numbers behind “most widely used”
scikit-learn’s case isn’t built on being the fastest — it’s built on how many people already depend on it, and how long it took to earn that without rushing a 1.0 label.
Where it’s actually used
Three-word version: not deep learning. scikit-learn’s own published case studies span finance, retail, and beyond, and the pattern in all of them is the same — structured, tabular data and a need for something fast, interpretable, and easy to put into production:
- Finance and insurance: AXA uses it to speed up car-accident compensation and detect insurance fraud; Zopa for credit risk modeling and loan pricing; J.P. Morgan across classification tasks in financial decision-making.
- Retail and e-commerce: Booking.com uses it for hotel recommendation, fraudulent-reservation detection, and customer-support workforce scheduling.
- Everywhere a baseline is needed first: before reaching for a neural network, a scikit-learn model — often a gradient-boosted tree or a plain logistic regression — is frequently how a team establishes whether the problem needs deep learning at all.
scikit-learn vs. Keras, PyTorch, and TensorFlow
The comparison here isn’t really “which is best” — it’s “which problem are you solving”:
- Data shape. scikit-learn is built for structured, tabular data: rows and columns, the kind that lives in a spreadsheet or a SQL table. Keras, PyTorch, and TensorFlow are built for unstructured data at scale — images, audio, text, video — where deep learning’s advantage over classical methods is largest.
- What “training” means. A
RandomForestClassifier.fit()call finishes in seconds to minutes on a laptop CPU. Training a neural network from the other three guides in this series is usually a GPU job measured in hours. - They’re used together, not instead of each other. A typical real project uses scikit-learn for data cleaning, feature engineering, and a baseline model, and only reaches for Keras, PyTorch, or TensorFlow once tabular methods have been tried and a case for deep learning is clear.
Frequently asked questions
Who created scikit-learn, and when?
David Cournapeau, as a Google Summer of Code project in June 2007, originally named scikits.learn. INRIA took over project leadership in 2010 and published the first public release that January.
Is scikit-learn the same as SciPy?
No, but they’re related: scikit-learn began as a third-party extension to SciPy and still depends on it and NumPy for numerical computation. SciPy is a general scientific-computing library; scikit-learn is specifically for machine learning.
Does scikit-learn support GPUs?
Partially, and only recently. Since version 1.8 (December 2025), estimators that support the Array API standard can run on a GPU if you pass them a GPU-backed PyTorch tensor or CuPy array — but this is experimental and only covers specific estimators so far, not the whole library.
Is scikit-learn used for deep learning?
No. It has no neural network training capability comparable to Keras, PyTorch, or TensorFlow. It’s used for classical machine learning — regression, classification, clustering — and for the data preprocessing that often precedes deep learning work.
Why is scikit-learn so widely used if it’s not deep learning?
Because most real-world machine learning problems are tabular, not images or text, and scikit-learn’s consistent fit()/predict()/transform() API makes trying and comparing algorithms fast. A 2022 Kaggle survey of nearly 24,000 developers found it to be the most widely used ML framework overall.