- New York
- https://minwu-ai.github.io/
Pinned Loading
-
llm_red_teaming
llm_red_teaming PublicA modular, extensible toolkit for red teaming large language models and NLP systems — covering adversarial text attacks (character, word, sentence, semantic), jailbreak evaluation via JailbreakBenc…
Python 3
-
ml_validation_framework
ml_validation_framework PublicInteractive ML model validation framework: explainability, weak spots, robustness, fairness — widget-driven, no-code demo
Python 1
-
multi_agent_otel_eval
multi_agent_otel_eval PublicEnterprise-grade GenAI observability and evaluation framework using OpenTelemetry conventions, with multi-agent orchestration tested on the Mind2Web benchmark.
Jupyter Notebook 1
-
rag_eval_framework
rag_eval_framework PublicProvider-agnostic RAG evaluation framework benchmarked on HotpotQA — 13 metrics, multi-prompt comparison, failure diagnosis, auto-generated audit report
Jupyter Notebook 1
-
genai_capability_bench
genai_capability_bench PublicA modular benchmark suite for evaluating GenAI capabilities across answer accuracy, truthfulness, instruction following, reasoning, RAG, tool use, and agentic task completion.
Jupyter Notebook
If the problem persists, check the GitHub status page or contact support.