Skip to content
View minw0607's full-sized avatar

Block or report minw0607

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. llm_red_teaming llm_red_teaming Public

    A modular, extensible toolkit for red teaming large language models and NLP systems — covering adversarial text attacks (character, word, sentence, semantic), jailbreak evaluation via JailbreakBenc…

    Python 3

  2. ml_validation_framework ml_validation_framework Public

    Interactive ML model validation framework: explainability, weak spots, robustness, fairness — widget-driven, no-code demo

    Python 1

  3. multi_agent_otel_eval multi_agent_otel_eval Public

    Enterprise-grade GenAI observability and evaluation framework using OpenTelemetry conventions, with multi-agent orchestration tested on the Mind2Web benchmark.

    Jupyter Notebook 1

  4. rag_eval_framework rag_eval_framework Public

    Provider-agnostic RAG evaluation framework benchmarked on HotpotQA — 13 metrics, multi-prompt comparison, failure diagnosis, auto-generated audit report

    Jupyter Notebook 1

  5. genai_capability_bench genai_capability_bench Public

    A modular benchmark suite for evaluating GenAI capabilities across answer accuracy, truthfulness, instruction following, reasoning, RAG, tool use, and agentic task completion.

    Jupyter Notebook

  6. Regulus Regulus Public

    An AI governance standards lookup powered by RAG and knowledge graphs. Submit an issue or observation, and Regulus retrieves applicable risks, regulatory standards, and cross-referenced guidance ac…

    Jupyter Notebook 1