Skip to content

Repository files navigation

OPFython

PyPI CI Documentation DOI License

OPFython is a Python implementation of the Optimum-Path Forest family of classifiers. It provides supervised, semi-supervised, unsupervised, and KNN-supervised models backed by NumPy and Numba.

This implementation follows LibOPF. Please cite the original LibOPF authors as well as OPFython when using it in research.

OPFython requires Python 3.11 or newer.

Installation

Install with pip:

python -m pip install opfython

Or add it to a UV-managed project:

uv add opfython

For an editable source checkout:

python -m pip install -e .

Quick start

import numpy as np

from opfython.models import SupervisedOPF

X_train = np.asarray([[0.0, 0.0], [0.1, 0.2], [1.0, 1.0], [1.1, 0.9]])
Y_train = np.asarray([0, 0, 1, 1])
X_test = np.asarray([[0.05, 0.1], [1.05, 1.0]])

classifier = SupervisedOPF()
classifier.fit(X_train, Y_train)
predictions = classifier.predict(X_test)

Labels must be zero-based and sequential. Pre-computed distance matrices can be supplied through each classifier's pre_computed_distance constructor argument.

Features, labels, and optional sample indexes must have matching sample counts; mismatches raise opfython.utils.exception.SizeError. Supervised and semi-supervised models also support a single labeled class. KNN models require 1 <= k < number of training samples.

When splitting a pre-computed matrix's dataset, retain its original sample indexes with stream.splitter.split_with_index and pass the corresponding indexes to fit and predict. The matrix must cover those indexes, not merely match the training subset's size. See the pre-computed distance example.

Classifiers

Class Purpose
SupervisedOPF Complete-graph supervised classification
KNNSupervisedOPF Supervised classification with learned KNN adjacency
SemiSupervisedOPF Learning from labeled and unlabeled samples
UnsupervisedOPF Density-based clustering and label propagation

The package also includes 47 distance metrics, random generators, OPF evaluation measures, dataset loaders and splitters, package logging and exception helpers, and converters for LibOPF binary datasets.

Clustering purity is independent of cluster numbering and supports more clusters than true classes. Evaluation measures require one prediction for each true label and handle class IDs missing from an evaluation split.

See the documentation and the examples/applications directory for complete workflows.

The usage guide describes array ownership, model lifecycle, index mapping, failure behavior, and trusted model persistence.

Development

Follow CONVENTIONS.md for the cpmux-derived Python and Google-style documentation rules. OPFython retains its public API, Apache license, and Python 3.11 support while using the compatible modern typing syntax.

uv sync --all-groups
uv run pytest
uv run pre-commit run --all-files
uv run --group docs sphinx-build -W -b html docs docs/_build/html
uv run --group docs sphinx-build -W -b doctest docs docs/_build/doctest
uv build --no-sources

Citation

@article{rosa2021simpa,
    title = {OPFython: A Python implementation for Optimum-Path Forest},
    author = {Gustavo H. {de Rosa} and Joao P. Papa},
    journal = {Software Impacts},
    pages = {100113},
    year = {2021},
    issn = {2665-9638},
    doi = {https://doi.org/10.1016/j.simpa.2021.100113}
}
@misc{rosa2021speedup,
    title = {Speeding Up OPFython with Numba},
    author = {Gustavo H. de Rosa and Joao Paulo Papa},
    year = {2021},
    eprint = {2106.11828},
    archivePrefix = {arXiv},
    primaryClass = {cs.LG}
}

OPFython is licensed under the Apache License 2.0.

About

🌳 A Python-inspired implementation of the Optimum-Path Forest classifier.

Topics

Resources

Code of conduct

Stars

38 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages