Enable AI inferencing on z/OS
Use the zopen package manager (QuickStart Guide) to install:
zopen install llamacpp- Clone the repository:
git clone https://github.com/zopencommunity/llamacppport.git
cd llamacppport- Build using zopen:
zopen build -vvSee the zopen porting guide for more details.
z/OS is big-endian. GGUF files published on Hugging Face are little-endian, so their tensor data has to be swapped before it can be used.
This port does that automatically: llama_model_loader detects the mismatch and
byteswaps each tensor as it is read, so a stock little-endian GGUF loads and runs
correctly with no preparation. Two things follow from it:
- the swap happens on every load, which makes loading a large model slower;
- mmap is turned off for that model, because the mapping is read-only and backed by the file, so it cannot be swapped in place. A warning is logged.
If you load the same model repeatedly, converting it once to big-endian avoids both costs:
python3 gguf-py/gguf/scripts/gguf_convert_endian.py model.gguf bigThe conversion rewrites the file in place, so keep a backup. A model that is already big-endian is loaded directly, with no swapping and with mmap available.
Note that the converter handles far fewer quantisations than the loader does: it
knows Q4_0, Q8_0, Q4_K, Q6_K, MXFP4 and NVFP4, and stops with Cannot handle type on anything else. The loader covers roughly 30 types, so several common
mixed quantisations - a Q4_K_M model, for instance, which also carries Q5_K
tensors - can only be run through the automatic swap. Conversion is an
optimisation where it is available, not a prerequisite.
Run with -v to see which path was taken; the loader logs
model endianness differs from the host, byteswapping tensor data when it swaps.
The zDNN backend uses the on-chip AI accelerator found on z16 (Telum I) and later. Detection is entirely at run time, so a single binary runs on z15 and below - the backend simply reports no devices there and llama.cpp schedules onto CPU/BLAS instead.
Set GGML_DISABLE_ZDNN=1 in the environment to turn the accelerator off on
hardware that has it, which is useful for A/B comparisons.
While building if an error is encountered in the ggml-cpu.cpp file (perhaps related to pthread), run zopen upgrade zoslib -y and try building again.
Contributions are welcome! Please follow the zopen contribution guidelines.