A bash script for creating, listing and extracting X4 Foundations cat/dat archive pairs using only standard Unix utilities (awk, find, sort, stat, md5sum, xargs, paste).
Catalogs written by this script are byte-for-byte identical to those produced by Egosoft's XRCatTool.exe 1.10, so it can replace the Windows-only packing step of a mod build on Linux and macOS.
Clone the repository and make the script executable:
git clone https://github.com/chemodun/X4-LinuxCatTool.git
cd X4-LinuxCatTool
chmod +x x4_cat_tool.shOr download just the script directly:
curl -O https://raw.githubusercontent.com/chemodun/X4-LinuxCatTool/main/x4_cat_tool.sh
chmod +x x4_cat_tool.shOptionally, copy it somewhere on your PATH (e.g. ~/.local/bin/) so you can run it from anywhere:
cp x4_cat_tool.sh ~/.local/bin/x4_cat_tool.shX4 Foundations stores its game assets in paired catalog/data files:
.cat- plain-text catalog; each line describes one file:<filepath> <size_bytes> <unix_timestamp> <md5_32hex>.dat- binary blob containing all files concatenated in the order listed in the.cat.
Files are split across multiple numbered catalog pairs (01.cat/01.dat, 02.cat/02.dat, …). Higher-numbered catalogs have higher priority: the same path in 07.cat overrides 01.cat. A size-0 entry in a higher catalog is a deletion marker - the file is treated as absent.
_sig.cat signature catalogs are automatically skipped.
./x4_cat_tool.sh [OPTIONS] <source> <command> <path_or_mask> [<dest_dir>]
./x4_cat_tool.sh [OPTIONS] <source_dir> c <out.cat>
| Argument | Description |
|---|---|
source |
Folder containing cat/dat files, or a single .cat file. For c: the folder to pack |
command |
x - extract ls - list c - create |
path_or_mask |
Path prefix or glob mask (see Filtering). For c: the output catalog name, which must end in .cat |
dest_dir |
Output directory (required for x, not used for ls and c) |
| Flag | Applies to | Description |
|---|---|---|
-f |
x c |
Force overwrite of already-existing output files or of an existing catalog |
-n |
x |
Skip MD5 hash verification after extraction |
-s |
x |
Strip the filter path prefix from output paths (see Strip prefix) |
-j [N] |
x |
Extract with N parallel workers. Without N, one per CPU core minus one (see Parallel extraction) |
-i <dir> |
c |
Additional input folder; repeatable. Files from later folders override same-named files from earlier ones |
-I <re> |
c |
Include only paths matching this regex; repeatable, OR-ed together |
-E <re> |
c |
Exclude paths matching this regex; repeatable, applied after -I |
-a |
c |
Append to an existing catalog instead of replacing it |
-v |
all | Verbose output (shows offset, size, dat file, hash result; for c lists every packed entry) |
-h |
all | Show help |
Prints a table of matching catalog entries to stdout. No files are written.
./x4_cat_tool.sh /game/x4 ls assets/textures
./x4_cat_tool.sh /game/x4 ls "libraries/*.xml"
./x4_cat_tool.sh /game/x4 ls t/Output (folder/multi-cat mode) includes a CAT column showing which catalog last defined each file:
SIZE DATE CAT PATH
------------ ------------------- ------------ ----
123456 2025-08-19 17:21:23 03.cat assets/textures/foo.dds
[deleted] 2025-09-01 10:00:00 07.cat assets/textures/old.dds
In single-cat mode the CAT column is omitted.
Extracts matching files into dest_dir, preserving the catalog path structure by default.
./x4_cat_tool.sh /game/x4 x assets/textures /tmp/out
./x4_cat_tool.sh /game/x4/01.cat x "libraries/*.xml" /tmp/out
./x4_cat_tool.sh -f /game/x4 x maps /tmp/outEach extracted file is logged:
Extract: assets/textures/ui/foo.gz
Hash verification runs automatically after each extraction unless -n is given. A warning is printed on mismatch but extraction continues.
With -s, the path_or_mask prefix is stripped from the output path so files land directly in dest_dir without the leading path hierarchy:
./x4_cat_tool.sh -s /game/x4 x assets/textures/ui/player_info /tmp/out
# → /tmp/out/icon_logbook_alerts.gz (not /tmp/out/assets/textures/ui/player_info/...)The log shows both paths when stripping is active:
Extract: assets/textures/ui/player_info/icon_logbook_alerts.gz -> icon_logbook_alerts.gz
Extraction is sequential unless -j is given. -j splits it across several worker processes, and the count is optional:
./x4_cat_tool.sh -j /game/x4 x '*' /tmp/out # auto: one worker per core, minus one
./x4_cat_tool.sh -j4 /game/x4 x '*' /tmp/out # exactly 4 workers
./x4_cat_tool.sh -j1 /game/x4 x '*' /tmp/out # sequential, the defaultWithout a number, the tool uses one worker per CPU core minus one, so the machine stays usable while a large extraction runs. The core count comes from nproc, getconf _NPROCESSORS_ONLN, sysctl -n hw.ncpu or NUMBER_OF_PROCESSORS, whichever exists first; if none does it falls back to a single worker.
A number you pass yourself wins, but it is capped at the core count - asking for more workers than there are cores only adds contention. The cap is reported when it applies:
Note : -j 99 capped at 12 (CPU cores)
The number of workers actually started is printed before extraction begins, and can be lower than requested when there are fewer files than workers:
Workers : 11
Each worker takes one contiguous slice of the file list, sized so every worker gets roughly the same amount of work - bytes plus a flat per-file charge, since extraction is process-bound and a run of tiny files is real work too. Each slice is hash-verified by the worker that wrote it, so verification is parallel as well, and the mismatch total across all workers is reported at the end:
Hash mismatches : 0
The one visible cost is ordering: with more than one worker the per-file Extract: lines from different workers interleave, so they no longer appear in catalog order. The extracted files themselves are identical either way. If a worker dies, the run stops with a non-zero exit status naming how many failed.
Packs a folder into a cat/dat pair. The .dat name is derived from the .cat name, so ext_01.cat writes ext_01.dat beside it.
./x4_cat_tool.sh ./my_mod c ext_01.cat
./x4_cat_tool.sh -v ./my_mod c /tmp/build/ext_01.catAn existing catalog is never replaced silently - pass -f to overwrite it or -a to append to it. This is deliberately stricter than XRCatTool.exe, which overwrites without asking.
Multiple input folders are merged with -i, and later folders win. This is how you overlay a patch tree on top of a base tree:
./x4_cat_tool.sh ./base -i ./overrides c ext_01.cat-I includes and -E excludes take regular expressions, not globs, matched exactly the way XRCatTool.exe matches its -include / -exclude: as a substring search against the lowercased path. So -E "content.xml" drops that file wherever it sits, while -E "^assets/" only anchors at the start. Repeating -I ORs the patterns; -E is applied afterwards to whatever survived.
This mirrors a typical two-catalog mod build, where the substitution catalog and the extension catalog are cut from one source tree:
./x4_cat_tool.sh -I "ego_debuglog/ui.xml" ./my_mod c subst_01.cat
./x4_cat_tool.sh -E "ego_debuglog/ui.xml" -E "content.xml" ./my_mod c ext_01.catNothing is excluded by default, so packing a working copy directly will pull in .git and other development files. Filter them out explicitly:
./x4_cat_tool.sh -E "^\.git" -E "^docs/" -E "\.bat$" ./my_mod c ext_01.cat-a appends the new entries to an existing pair instead of rewriting it:
./x4_cat_tool.sh -a ./more_files c ext_01.catThe appended block is sorted within itself and added at the end, exactly as XRCatTool.exe -append does. The catalog as a whole is therefore no longer globally sorted, and duplicate paths are not merged away - the later entry simply wins at load time. If -a is given but the catalog does not exist yet, it is created.
Applies to x and ls. The c command filters with regexes instead, via -I / -E - see Filtering what gets packed.
| Pattern | Behaviour |
|---|---|
* |
Everything - the whole catalog. Use this to list or extract a complete archive |
assets/textures |
Prefix match - all files whose path starts with assets/textures |
assets/textures/ |
Same - trailing slash is stripped before matching |
libraries/*.xml |
Glob with / - matched against the full catalog path |
0001*.xml |
Glob without / - matched against the filename only (e.g. matches t/0001-l044.xml) |
An empty mask ("") matches nothing rather than everything, so pass * when you want the whole catalog:
./x4_cat_tool.sh /game/x4/01.cat x '*' /tmp/outCat files are sorted by name ascending. Each subsequent file overrides earlier entries for the same path:
01.cat assets/foo.bin 500000 ... ← earlier version
07.cat assets/foo.bin 0 ... ← deletion marker → file is NOT extracted
Only the specified .cat/.dat pair is used. Size-0 entries produce no output.
# List all files under the 't' directory
./x4_cat_tool.sh '/c/Program Files (x86)/X4 Foundations' ls t/
# List all XML files matching a name pattern across all catalogs
./x4_cat_tool.sh '/c/Program Files (x86)/X4 Foundations' ls '0001*.xml'
# Extract all libraries
./x4_cat_tool.sh '/c/Program Files (x86)/X4 Foundations' x libraries/ /tmp/x4out
# Extract a texture folder, stripping the leading path
./x4_cat_tool.sh -s '/c/Program Files (x86)/X4 Foundations' x \
assets/textures/ui/player_info /tmp/icons
# Extract from a specific catalog only (no priority merging)
./x4_cat_tool.sh '/c/Program Files (x86)/X4 Foundations/01.cat' x maps/ /tmp/maps
# Force re-extract, skip hash check, verbose
./x4_cat_tool.sh -fnv '/c/Program Files (x86)/X4 Foundations' x aiscripts/ /tmp/ai
# Extract an entire catalog
./x4_cat_tool.sh '/c/Program Files (x86)/X4 Foundations/01.cat' x '*' /tmp/all
# Extract an entire catalog with one worker per core (minus one)
./x4_cat_tool.sh -j '/c/Program Files (x86)/X4 Foundations/01.cat' x '*' /tmp/all
# Pack a mod folder into a catalog pair
./x4_cat_tool.sh ./my_mod c ext_01.cat
# Pack everything except the assets tree, listing each entry
./x4_cat_tool.sh -v -E '^assets/' ./my_mod c ext_01.catThese rules were confirmed by packing test trees with XRCatTool.exe 1.10 and comparing the bytes. They matter if you ever hand-write or post-process a catalog:
- The
.catuses LF line endings with no header and no BOM, and ends with a trailing newline. - Paths keep their original case and always use forward slashes.
- Entries are sorted by the lowercased path in C byte order. So
UPPER.TXTsorts afterstray.dat, not before it, and a plainsorton the raw paths gives the wrong order. - The timestamp is the source file's mtime, not the time of packing.
- A zero-byte file is written with 32 zeros as its hash. The md5 of an empty file (
d41d8cd9…) never appears. - The
.datis a plain concatenation of the file contents in catalog order - no header, no padding, no alignment. Offsets are implied by the running sum of the sizes.
A deletion marker, which the game reads as "this file is absent", is an entry with size 0 and timestamp 0. c does not generate those; it is what XRCatTool.exe -diff emits.
- Bash 4.0+ (associative arrays,
${var,,}) awk,find,sort,tail,head,date,mkdir,dirnamemd5sum(Linux) ormd5(macOS)- For
cadditionally:stat,paste,tr,xargs,mktemp,sed,cat,wc
Available out of the box on any Linux distribution and macOS. Both GNU (stat -c) and BSD/macOS (stat -f) syntax are detected automatically.
-j needs nothing beyond bash itself. Core detection tries nproc, getconf, sysctl and NUMBER_OF_PROCESSORS in that order, and falls back to a single worker if none of them answers.
Source paths containing tab or newline characters are not supported, as the pipelines are tab- and line-delimited. Neither is representable in an X4 mod.
The entire catalog merge and filter step runs inside a single awk invocation. This makes parsing 450,000+ entries across 9 catalog files feasible in a few seconds.
c works the same way. The file scan, filtering, sorting and metadata collection are pipelines rather than per-file shell loops, and hashing runs as one md5sum process per xargs batch instead of one per file. Packing a 1000-file, 20 MB mod takes a few seconds.
x is process-bound rather than I/O-bound, since tail -c +N and dd both seek straight to the offset - a file at the end of a 20 GB archive costs no more to extract than one at the start. Extraction therefore runs one process per file rather than the six or seven a naive loop needs: output paths are resolved with shell builtins instead of dirname, every output directory is created in one batched mkdir, the byte range is copied by a single dd where GNU dd is available (falling back to tail | head on macOS and BSD), and hashes are verified in xargs batches after the fact instead of one md5sum per file.
Because that remaining process per file is the floor, -j is what cuts into it - the workers spend most of their time in dd and md5sum, which run concurrently. Extracting a 1039-file, 20 MB catalog on a 12-core machine, with hash verification on:
-j 1 61s
-j 2 33s
-j 4 23s
-j 11 17s
Returns flatten out well before the core count is reached, so there is little point forcing a high -j by hand - and none at all above the cores available, which is why the tool caps it there. The measurements above are from Git Bash on Windows, where creating a process is unusually expensive; on Linux the sequential baseline is far lower and the parallel gain correspondingly smaller.
- Added:
-j [N]- extract with several worker processes at once. The count is optional: a bare-juses one worker per CPU core minus one, so the machine stays usable, while-j 4sets it explicitly. A hand-picked count is capped at the number of cores, and-j 1keeps the sequential path. Each worker writes and hash-verifies its own contiguous slice of the file list, sized by bytes plus a flat per-file charge so the slices cost about the same. Measured on a 1039-file catalog with verification on: 61s sequential, 23s at-j 4, 17s at-j 11.- The number of hash mismatches is now reported in the extraction summary instead of only appearing as individual warnings.
- Changed:
- Extraction is substantially faster. A file now costs one process instead of six or seven - output paths are resolved with shell builtins, every output directory is created in a single batched
mkdir, the byte range is copied by oneddwhere GNUddis available, and hashes are verified in batches afterwards rather than onemd5sumper file. Measured 3.4x faster with verification and 2.6x with-n; a full 1000-file catalog that previously ran past a two-minute timeout now completes in about a minute. Gains are largest where process creation is expensive.
- Extraction is substantially faster. A file now costs one process instead of six or seven - output paths are resolved with shell builtins, every output directory is created in a single batched
- Fixed:
- With
-s, the check that skips already-extracted files tested the unstripped path, so it never matched and every file was re-extracted on each run. It now checks the real output path.
- With
- Added:
ccommand - create acat/datpair from a folder. Output is byte-for-byte identical toXRCatTool.exe1.10, verified across nine cases including a 1000-file mod.-i <dir>- additional input folders, repeatable, with files from later folders overriding earlier ones.-I <re>and-E <re>- include and exclude filters forc, using the same regex substring-search semantics asXRCatTool.exe's-include/-exclude. Both are repeatable;-Ipatterns are OR-ed and-Eis applied afterwards.-a- append to an existing catalog rather than replacing it.-V- print the tool version and exit.
- Changed:
-fnow also guardsc: an existing catalog is never overwritten without it. This is deliberately stricter thanXRCatTool.exe, which overwrites silently.
- Documentation:
- The
*mask is now documented, including that an empty mask matches nothing rather than everything. - New section describing the catalog format written by
c- the case-folded sort order, the zero hash used for empty files, and the size-0 / timestamp-0 shape of a deletion marker.
- The
- Added:
lscommand - list catalog entries with size, date and originating catalog.xcommand - extract entries, with MD5 verification (-nto skip) and-sto strip the filter prefix from output paths.- Folder mode with catalog priority merging and deletion markers, and single-catalog mode.
- Prefix, full-path glob and filename-only glob filtering.