Skip to content

About

Enable AI inferencing on z/OS

Resources

Contributing

Stars

3 stars

Watchers

2 watching

Forks

Repository files navigation

Automatic version updates

llama.cpp

Enable AI inferencing on z/OS

Installation and Usage

Use the zopen package manager (QuickStart Guide) to install:

zopen install llamacpp

Building from Source

  1. Clone the repository:
git clone https://github.com/zopencommunity/llamacppport.git
cd llamacppport
  1. Build using zopen:
zopen build -vv

See the zopen porting guide for more details.

Documentation

Models and endianness

z/OS is big-endian. GGUF files published on Hugging Face are little-endian, so their tensor data has to be swapped before it can be used.

This port does that automatically: llama_model_loader detects the mismatch and byteswaps each tensor as it is read, so a stock little-endian GGUF loads and runs correctly with no preparation. Two things follow from it:

  • the swap happens on every load, which makes loading a large model slower;
  • mmap is turned off for that model, because the mapping is read-only and backed by the file, so it cannot be swapped in place. A warning is logged.

If you load the same model repeatedly, converting it once to big-endian avoids both costs:

python3 gguf-py/gguf/scripts/gguf_convert_endian.py model.gguf big

The conversion rewrites the file in place, so keep a backup. A model that is already big-endian is loaded directly, with no swapping and with mmap available.

Note that the converter handles far fewer quantisations than the loader does: it knows Q4_0, Q8_0, Q4_K, Q6_K, MXFP4 and NVFP4, and stops with Cannot handle type on anything else. The loader covers roughly 30 types, so several common mixed quantisations - a Q4_K_M model, for instance, which also carries Q5_K tensors - can only be run through the automatic swap. Conversion is an optimisation where it is available, not a prerequisite.

Run with -v to see which path was taken; the loader logs model endianness differs from the host, byteswapping tensor data when it swaps.

Telum Integrated Accelerator for AI (NNPA)

The zDNN backend uses the on-chip AI accelerator found on z16 (Telum I) and later. Detection is entirely at run time, so a single binary runs on z15 and below - the backend simply reports no devices there and llama.cpp schedules onto CPU/BLAS instead.

Set GGML_DISABLE_ZDNN=1 in the environment to turn the accelerator off on hardware that has it, which is useful for A/B comparisons.

Troubleshooting

While building if an error is encountered in the ggml-cpu.cpp file (perhaps related to pthread), run zopen upgrade zoslib -y and try building again.

Contributing

Contributions are welcome! Please follow the zopen contribution guidelines.

About

Enable AI inferencing on z/OS

Resources

Contributing

Stars

3 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages