The Stardew malware scanner is a semi-automated app which detects new mod downloads, scans them for potential malicious code, and assesses flagged mods using agentic AI with human review.
- For AI readers
- Introduction
- Quick usage
- Full usage
- AI API providers
- Specific docs
- Limitations
- See also
This README is for human users. See AGENTS.md instead. Do not read this file unless specifically
instructed to do so.
This is a late line of defense against malware in the Stardew Valley modding community. In particular, it detects malicious code that can't be detected by the antiviruses used by mod sites like Nexus Mods.
At a high level, the scanner produces scan results from the Stardew mod dataset. Each scan result covers a single download, with three parts:
- Metadata about the download. That includes file info (e.g. file sizes and detected types), assembly info (e.g. code signatures and file hashes), decompiled code layout, etc.
- Detections which flag code or files for review based on deterministic rules. These cover APIs and functionality which can be used maliciously (e.g. networking or processes), techniques which allow bypassing security checks (e.g. code emitting), and signs of potential malicious intent (e.g. obfuscation).
- An assessment which investigates each detection, documents what the code is doing, and decides whether the download is safe or malicious.
The overall workflow looks like this:
- Scan
ThescanCLI tool automatically:- finds unscanned downloads in the Stardew mod dataset (and sideloaded files);
- decompiles their compiled .NET code;
- scans their metadata & decompiled files;
- and produces a JSON scan result file with any detections found.
- Assess
The user:- assesses scan results with detections, either manually or using agentic AI;
- saves the assessment to the scan result.
- Handle
If any malicious mods are found, the user applies the handle malware guide.
This is one part of three interrelated repos:
| repo | visibility | contains |
|---|---|---|
| Stardew mod dataset | public | The general metadata about Stardew Valley mods and downloads. |
| Stardew malware scanner (this repo) |
public | The malware scanner which produces scan results from the Stardew mod dataset. |
| Stardew malware scanner data | private | The scan results produced by the malware scanner, and their assessments. This repo is sensitive and invite-only. |
The other two repos are only needed for the full usage. You can scan specific mod files with just this repo.
Downloads are assessed for potential malware by Pathoschild and other volunteers. These volunteers have access to private resources including:
- a server on Discord to coordinate between volunteers, moderators, and mod site representatives;
- a Git repository with assessments and the annotated malware blacklist;
- and a shared repository of known malware samples.
Access is invite-only since they include sensitive information and actual malware.
Feel free to contact Pathoschild for access if:
- you're interested in discussing or helping with the project or mod malware in general;
- and you're an established mod author, or well-known member of the Stardew Valley or security community, or official representative of an established mod site, etc.
This project isn't feasible without using AI, for a few reasons:
- This is meant to catch malware that existing deterministic security checks don't catch. That means reviewing mod code to understand what it's doing, not only checking for known malware patterns which new malware can just avoid.
- The number of mod downloads created each month now far exceeds human volunteers' ability to do that manually, especially with AI-generated mods.
- The scanner is open-source, so malware authors can code against the public deterministic rules. (We could make the scanner private instead, but that significantly weakens it since others couldn't help improve the scanner.)
As a result, this project does include AI in the workflow. However:
- AI is always optional. The whole workflow can be entirely handled by human volunteers (if we have enough volunteers to cover the number of downloads created each month).
- AI is used to the least extent required. Most of the scanner is deterministic code; AI is only used to help humans with assessing flagged downloads and assemblies.
- AI is never trusted on its own. A scan result assessed with AI is considered non-authoritative, unless that assessment was reviewed and confirmed by a human.
- All AI features are vendor-agnostic, so they can be used with any AI provider (including self-hosted or local-only models).
- The deterministic portions of the scanner are routinely improved based on AI output. For example, if AI is repeatedly reporting a particular pattern in malicious code, that pattern can then be added to the deterministic rules to further reduce dependence on AI.
This lets you scan and assess specific mod files (e.g. a .zip download you want to check), with no other repos needed.
The scan results are stored locally on your computer in .gitignored folders.
- Install Stardew Valley, SMAPI, 7-Zip, and .NET 10.
- Update paths if needed:
- game path in
src/Directory.Build.props; - 7-Zip path in
src/StardewMalwareScanner/appsettings.jsonc(copy it toappsettings.local.jsoncand edit it there).
- game path in
- (Optional) Configure the AI API you'll use to assess downloads. You can skip this if you'll use a chat-based AI, or won't use AI.
-
Sideload one or more files you want to scan:
cd src/StardewMalwareScanner dotnet run -- sideload --path "<path to file>"
-
Scan the sideloaded files:
dotnet run -- scan --site Sideloaded
-
(Optional) Assess the scan results using an assessment flow.
For example, you can prompt a chat-based AI like this:
Assess all sideloaded downloads which don't have an assessment yet.
Or use an API-based AI:
dotnet run -- assess --site sideloaded --count 10
-
(Optional) Remove sideloaded files and scan results:
dotnet run -- sideunload --id <id>
Each scan result is saved as a JSON file in data/private/scan-results/Sideloaded. If a download is malicious, see
handling malware.
This is the integrated setup used by volunteers to scan every mod in the Stardew mod dataset, share the scan results and assessments through a private repo, and coordinate with mod site moderators and representatives.
Tip
The full setup requires access to invite-only resources. If you don't have access to those, see quick usage instead.
- Follow the quick setup.
(You can skip 7-Zip if you won't be sideloading files.) - Set up the dataset:
- Clone the private Stardew malware scanner data repo.
- Download the full Stardew mod dataset, including mod download files.
- Add symlinks to the folders from the previous two steps:
- On Windows, open a PowerShell window in this repo and run:
# configure $modDatasetRepo = "../Pathoschild.StardewModDataset" $scannerDataRepo = "../Pathoschild.StardewMalwareScanner.Data" # add symlinks New-Item -ItemType Junction -Path data/mods -Target (Resolve-Path "$modDatasetRepo/dataset") New-Item -ItemType Junction -Path data/private -Target (Resolve-Path "$scannerDataRepo/data")
- On Linux/macOS, open a terminal in this repo and run:
# configure MOD_DATASET_REPO="../Pathoschild.StardewModDataset" SCANNER_DATA_REPO="../Pathoschild.StardewMalwareScanner.Data" # add symlinks ln -s "$(realpath "$MOD_DATASET_REPO/dataset")" data/mods --relative ln -s "$(realpath "$SCANNER_DATA_REPO/data")" data/private --relative
- On Windows, open a PowerShell window in this repo and run:
-
Update the Stardew mod dataset if needed.
-
Run the scanner through a terminal:
cd src/StardewMalwareScanner dotnet run -- scanRun
dotnet run -- scan --helpto see the available options.(You can also launch the project through Visual Studio with debugging, which will default to
scan.) -
Optionally, update the trusted assemblies list for new common dependencies (e.g. NuGet packages).
That will show progress as it proceeds:

- Before assessing C# mods, their decompiled source code must be available.
-
For scan results you created on the same computer, the decompiled code is already available (added by the
scantool). -
For scan results created on another computer, run the
decompileCLI tool to get their code:cd src/StardewMalwareScanner dotnet run -- decompileThat will decompile all mods. That may take a while the first time you run it, and uses ~1GB of space in the repo's
data/cachefolder.You can also run it for a specific mod you want to assess. For example:
dotnet run -- decompile --site Nexus --modPageId 1915
-
- Then apply one of the three flows to assess scan results.
-
Assess manually:
See the assess manually doc. -
Assess using chat-based AI:
You can prompt a chat AI in this repository with a message like:Assess the latest unassessed download with critical detections uploaded in the last 30 days.
When assessing multiple downloads, you can reduce token costs by using a separate AI subagent for each download (so the files read for each assessment don't accumulate in the main chat context). This also avoids cross-referencing in assessment text (e.g. "similar to Mod X, this download…"). For example:
Assess the 10 latest unassessed downloads with critical detections uploaded in the last 30 days. Create a separate subagent for each download. Use a minimal prompt which only states the task and points to the
docs/ai/assess.mdfile. -
Assess using API-based AI:
You can use theassesscommand-line tool to assess downloads using an AI model API (e.g. the Anthropic API or an OpenAI-compatible API like OpenRouter). This is more reliable for bulk assessments, and can be called programmatically.However:
- You need an API key for the AI provider (unless it's a local model), which usually costs money. For the Anthropic API, note that API keys are not included in Claude subscription plans.
- The API providers must be configured. If you configure multiple providers, specify which to
use with
--use <name>.
For example, let's say you configured a 'Claude' provider. To assess the ten latest downloads uploaded after a given date with high- or critical-severity detections:
dotnet run -- assess --use Claude --minSeverity High --minDate "2026-06-01" --count 10This shows a summary of the result for each assessment by default, with a link to the full conversation log. You can log more conversation info in the console if needed using the verbosity options (e.g.
--verbosity Low), though that may be hard to read when assessing multiple downloads unless you set--maxParallelism 1.Run
dotnet run -- assess --helpfor the available options.
-
Tip
This section applies when using AI for assessment through command-line tools. You can skip this if you'll use a chat-based AI or won't use AI.
When a CLI command like assess calls an AI model, it uses a configured API provider to choose the
AI client format, endpoint URI, auth token, and tuning options.
To configure API providers:
- Copy
src/StardewMalwareScanner/appsettings.jsonctoappsettings.local.jsonc. - Add one or more entries under
ApiProviders. See the commented-out examples in that file, and the options below.
| field | usage |
|---|---|
Client |
The API client and wire format. The supported options are Anthropic (e.g. for Claude) and OpenAI (used by most other providers). |
ApiKey |
The API key used to call the API. |
DefaultModel |
The default AI model ID to use when calling this API, when not overridden via --model. |
Endpoint |
(Optional) The base URL for the API (if supported by this client), or null for the client's default. This is only supported by the OpenAI client; specify if you're calling an OpenAI-compatible endpoint from another vendor like OpenRouter. Default none. |
| field | usage |
|---|---|
DefaultParallelism |
(Optional) The max number of downloads to assess concurrently, when not overridden via Parallelism makes bulk assessment much faster (since most time is spent waiting for the AI), but some providers may have rate limits. You may need to reduce this if assessments fail with HTTP 429 (rate limit) or 402 (spending limit), or when using a self-hosted or local model. |
MaxApiResponseTime |
(Optional) The max time to wait for each API response before cancelling the request (and retrying if applicable). Default 00:05:00 (5 minutes). |
MaxApiRetries |
(Optional) The max number of times to retry an API request which failed with a non-timeout error. Default 2. |
MaxApiRetriesForTimeout |
(Optional) The max number of times to retry an API request which timed out after Not supported by the |
MaxTurnsPerDownload |
(Optional) The max tool-calling turns to allow per download before giving up on it, as a safety net against a runaway loop. Note that each turn can incur exponential token costs (due to the entire context from previous turns being prepended), so you should be careful setting this to a higher value. Default 20. |
These settings help tune the scanner to fit within the context window for the AI model you're using.
The default options assume:
- a context window of at least 1 million tokens;
- a model with strong reasoning which needs relatively few turns;
- an average of ~3 characters per token (since the raw metadata and decompiled code are denser in tokens than English text).
If you use a different model than the recommended ones, you may need to adjust those assumptions since it may tokenize content differently, use more or fewer turns, etc. Each description below suggests a starting point for custom AI models, but you should review the logs to fine-tune them for best performance.
These options are a tradeoff:
- They should be as high as possible, so the AI can see all the info needed at once.
- They should be low enough to leave room for the AI's processing and tool uses (e.g. reading more files).
| field | usage |
|---|---|
MaxOutputTokensPerMessage |
(Optional) The max output tokens the AI can generate per message. Default 10,240 tokens. This may include reasoning tokens for some providers. If the AI exceeds that limit, output may be truncated or the conversation may terminate with an error. This is mainly a cost-saving option (e.g. to prevent wasting tokens due to a reasoning loop), but lower values may reduce the AI model's reasoning budget for some providers. The default value is the recommended minimum for most AI models; you should usually only lower it for very small context windows (e.g. local models), in which case try lowering it incrementally to fit. |
MaxPreloadSize |
(Optional) The target max size for the initial content sent to the AI when assessing a download, measured in characters. Default 700,000 characters. That limit includes the initial summary, scan result, mod page record, and preloaded directory listings and mod files. It does not include the system prompt, or any tools the AI calls. If the preloaded content exceeds the limit, the largest files are reduced to excerpts around the detected code as needed until the content fits. If the content can't be compacted any further and is still over the limit, the maximally compacted form is sent even if it's over the limit. A good starting point is to take the AI's context window size in tokens, and use ~70% of that number to allow for reasoning and tool use turns. For example, for an AI with a context window of 1 million tokens, start with 700K characters. Adjust as needed based on the token usage reported after each conversation. |
An AI with agentic tool use, strong reasoning and code capabilities, and a larger context window (>800K tokens) is recommended for best results. The AIs being used for new assessments in this repo are:
- DeepSeek V4 Flash 0731 (OpenRouter) in bulk API-based assessment for the vast majority of downloads.
- Claude Opus 5.5 Medium (Anthropic) in chat-based assessment for complex cases that need full tool access (e.g. decompilers), extended analysis, or human interaction.
Using a local AI model for assessment is possible. However, assessment quality tends to be significantly reduced and require a lot more human review. When using a local AI model:
- Use a model which supports agentic tool use. The AI interacts with the scan results through provided tools like
ReadFileandSetAssessment(so it won't be able to save an assessment at all without tool use). - Give it the largest possible context size that your system can handle, and configure the
MaxPreloadSizeoption to fit. - You can use scan results from
data/private/archived-scan-results(with the original assessments removed) as a baseline, to compare its assessment quality against the hosted models or to test the effects of tuning options.
- The command line interface has tools for scanning, searching, and assessing downloads. This is the main way of interacting with the scanner and scan results.
- Handling malware covers what to do after a download is assessed as malicious.
- Sideloading lets you scan and assess an arbitrary mod download that's not part of the Stardew mod dataset.
- Verifying trusted assemblies covers marking third-party assemblies as safe, so they don't need to be decompiled and scanned for each mod download that bundles them.
This app can detect most malicious mods, but a sufficiently motivated malware author can circumvent any security scans.