OPFython is a Python implementation of the Optimum-Path Forest family of classifiers. It provides supervised, semi-supervised, unsupervised, and KNN-supervised models backed by NumPy and Numba.
This implementation follows LibOPF. Please cite the original LibOPF authors as well as OPFython when using it in research.
OPFython requires Python 3.11 or newer.
Install with pip:
python -m pip install opfythonOr add it to a UV-managed project:
uv add opfythonFor an editable source checkout:
python -m pip install -e .import numpy as np
from opfython.models import SupervisedOPF
X_train = np.asarray([[0.0, 0.0], [0.1, 0.2], [1.0, 1.0], [1.1, 0.9]])
Y_train = np.asarray([0, 0, 1, 1])
X_test = np.asarray([[0.05, 0.1], [1.05, 1.0]])
classifier = SupervisedOPF()
classifier.fit(X_train, Y_train)
predictions = classifier.predict(X_test)Labels must be zero-based and sequential. Pre-computed distance matrices can
be supplied through each classifier's pre_computed_distance constructor
argument.
Features, labels, and optional sample indexes must have matching sample
counts; mismatches raise opfython.utils.exception.SizeError. Supervised
and semi-supervised models also support a single labeled class. KNN models
require 1 <= k < number of training samples.
When splitting a pre-computed matrix's dataset, retain its original sample
indexes with stream.splitter.split_with_index and pass the corresponding
indexes to fit and predict. The matrix must cover those indexes, not
merely match the training subset's size. See the
pre-computed distance example.
| Class | Purpose |
|---|---|
SupervisedOPF |
Complete-graph supervised classification |
KNNSupervisedOPF |
Supervised classification with learned KNN adjacency |
SemiSupervisedOPF |
Learning from labeled and unlabeled samples |
UnsupervisedOPF |
Density-based clustering and label propagation |
The package also includes 47 distance metrics, random generators, OPF evaluation measures, dataset loaders and splitters, package logging and exception helpers, and converters for LibOPF binary datasets.
Clustering purity is independent of cluster numbering and supports more clusters than true classes. Evaluation measures require one prediction for each true label and handle class IDs missing from an evaluation split.
See the documentation and the
examples/applications directory for complete
workflows.
The usage guide describes array ownership, model lifecycle, index mapping, failure behavior, and trusted model persistence.
Follow CONVENTIONS.md for the cpmux-derived Python and Google-style documentation rules. OPFython retains its public API, Apache license, and Python 3.11 support while using the compatible modern typing syntax.
uv sync --all-groups
uv run pytest
uv run pre-commit run --all-files
uv run --group docs sphinx-build -W -b html docs docs/_build/html
uv run --group docs sphinx-build -W -b doctest docs docs/_build/doctest
uv build --no-sources@article{rosa2021simpa,
title = {OPFython: A Python implementation for Optimum-Path Forest},
author = {Gustavo H. {de Rosa} and Joao P. Papa},
journal = {Software Impacts},
pages = {100113},
year = {2021},
issn = {2665-9638},
doi = {https://doi.org/10.1016/j.simpa.2021.100113}
}@misc{rosa2021speedup,
title = {Speeding Up OPFython with Numba},
author = {Gustavo H. de Rosa and Joao Paulo Papa},
year = {2021},
eprint = {2106.11828},
archivePrefix = {arXiv},
primaryClass = {cs.LG}
}OPFython is licensed under the Apache License 2.0.