Statistical evaluation of Gemma refusal robustness using direction ablation and official SORRY-Bench scoring.
-
Updated
Jun 25, 2026 - Python
Statistical evaluation of Gemma refusal robustness using direction ablation and official SORRY-Bench scoring.
Quantitative performance benchmarking of machine learning classifiers for phishing detection, utilizing precision, recall, and F1-score optimization on high-dimensional labeled datasets.
Classification models for detecting fake reviews and predicting software bugs. Includes implementations of decision trees, bagging, random forests, logistic regression, and Naive Bayes, with statistical evaluation using McNemar's test.
frontier-evals-harness is a lightweight framework for benchmarking frontier language models. It provides deterministic suite versioning, modular adapters, standardized scoring, and paired statistical comparisons with confidence intervals. Built for regression tracking and analysis, it enables reproducible evaluation without infrastructure.
To associate your repository with the statistical-evaluation topic, visit your repo's landing page and select "manage topics."