Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
133 commits
Select commit Hold shift + click to select a range
0666ad2
ci : target ROCm 7.14 for build and release (#25775)
superm1 Aug 10, 2026
689e227
opencl: transpose the K tile in local memory for FA prefill kernels (…
wanghqc Aug 10, 2026
030ebb5
Address review comment of PR 25532 (#26852)
gaugarg-nv Aug 10, 2026
84f7129
ggml-webgpu: fix CI errors from #25025 and #25262 (#26566)
yomaytk Aug 11, 2026
48d22e2
common/peg : suppress incomplete escape sequences (#26780)
aldehir Aug 11, 2026
14e78dd
model : fix SWA not being enabled for EXAONE 4.5 (#26848)
junmo-kim Aug 11, 2026
4801e3c
tests : disable backend sampler hip multi output (#26878)
jimw567 Aug 11, 2026
b3df572
tests : clean-up server test, use `tests.sh` in ci (#26886)
ggerganov Aug 11, 2026
153d324
llama: add default load-mode auto, which avoids mmap on iGPUs (#26081)
0cc4m Aug 11, 2026
9afff1b
tests : fix running server tests on windows (#26889)
ggerganov Aug 11, 2026
1138b85
model-conversion : use save_output_data for causual embeddings [no ci…
danbev Aug 11, 2026
7044859
ci: hip-quality-check: update vgpr spill ignore list (#26859)
IMbackK Aug 11, 2026
8d274dd
ui: fix context gauge for single-model usage (#25738)
intel00000 Aug 11, 2026
6e62ba5
mtmd: support pocket-tts (#26871)
ngxson Aug 11, 2026
cc078b4
Dflash support for nemotron-3.5 (#26905)
lnigam Aug 11, 2026
5d16e81
convert : keep quantization scales for nemotron --mtp export (#26903)
ynankani Aug 11, 2026
2468576
requirements: use stable torch packages on s390x (#26864)
nikwen Aug 11, 2026
38406d5
imatrix.cpp: Move finite check and only check touched experts (#26861)
bartowski1182 Aug 11, 2026
70dfba5
ci : add windows-rocm to check-release (#26897)
CISC Aug 11, 2026
ba360ef
chat : tighten bare function parsing for Qwen models (#26793)
aldehir Aug 11, 2026
f785fc9
spec : update speculative-simple (#26904)
ggerganov Aug 11, 2026
5988633
cuda : add warp-per-row wkv7 kernel for single-token decode (#26111)
123123213weqw Aug 11, 2026
ebb546b
CUDA: only disable CUDA graphs when mul_mat_id actually needs a strea…
grafail Aug 11, 2026
7b13a84
ci : add missing release check (#26923)
CISC Aug 11, 2026
0b1bad1
chat : fix muse-glimmer detection of tool calls after EOM (#26879)
ruanslv Aug 11, 2026
cb27fe9
opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (…
lhez Aug 12, 2026
6eff593
convert : handle per_layer_config in Gemma4 (transformers 5.15) (#26882)
pluvium27 Aug 12, 2026
55f453b
wavtokenizer-dec : bound posnet/convnext block_count against n_layer_…
oakkaya Aug 12, 2026
a7cd2f0
vulkan: add TQ2_0 (ternary) support (#25850)
michaeltrabalka-tech Aug 12, 2026
a4a4c51
tests : update speculative params (#26925)
ggerganov Aug 12, 2026
89e0aa6
opencl: default FA c8 cluster width to 16 on X1E (#26433)
wanghqc Aug 12, 2026
4dd1275
ui: add read_media tool (#25877)
parabelboi Aug 12, 2026
5d9e5ac
server : support slot save/restore with media inputs (#26640)
CHIPMUNK-T0T Aug 12, 2026
13fd0bb
cmake : add config version support (ggml/1582)
danbev Aug 12, 2026
af05a42
sync : ggml
ggerganov Aug 12, 2026
ece98b8
model : disallow integer dflash sliding_window_pattern (#26900)
CISC Aug 12, 2026
132753b
kleidiai: Add runtime feature detection mechanism for aarch64/kleidia…
JonathanC-ARM Aug 12, 2026
d8a8bea
gguf : harden loader against malformed tensor dims and metadata types…
harrison001 Aug 12, 2026
680a9ae
cmake : introduce semantic versioning (#26839)
danbev Aug 12, 2026
7a9ff95
disable rocm cache (#26962)
ggerganov Aug 12, 2026
9558fa4
ci : disable ubuntu-rocm (#26969)
ggerganov Aug 12, 2026
84e908c
ci: fix thread sanitizer + remove ccache (#26927)
netrunnereve Aug 12, 2026
8e7f22b
common: add system-level config file (#26118)
jcmdln Aug 12, 2026
e21152d
ui: Constants refactor (#26908)
ServeurpersoCom Aug 13, 2026
1f368f3
ggml : fix arm builds, unused var (#26991)
ggerganov Aug 13, 2026
a6040c9
refactor: Clean up UI types (#26909)
allozaur Aug 13, 2026
094e53d
ui: Stores architecture improvements (#26910)
allozaur Aug 13, 2026
f2efd64
ui: Move `styles/` to `$lib` scope (#26950)
allozaur Aug 13, 2026
d86c7d6
ui: Clean up contexts, remove prop drilling from Chat Form Actions (#…
allozaur Aug 13, 2026
e79e4bf
ggml-hip : remove -funsafe-math-optimizations (#26696)
jimw567 Aug 13, 2026
decaf50
server: refactor + correctness fixes for metrics (#26920)
ngxson Aug 13, 2026
d415e65
sycl : enhance concat to support Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 (#26800)
arthw Aug 13, 2026
8efbf65
sycl : Add DMMV ESIMD Q3_K kernel (#26251)
malsbat Aug 13, 2026
1ee1cd9
sycl: fuse UNARY(silu|sigmoid|softplus) + MUL (#26411)
Titaniumtown Aug 13, 2026
154d57a
sycl: remove separate fp32 type promotion in gemm non-oneDNN path (#2…
icfaust Aug 13, 2026
eeae28b
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion…
jinzihao Aug 13, 2026
0d0bfcd
spec: enable backend sampling for both dflash & dspark (#26958)
ruixiang63 Aug 13, 2026
f65e568
common : auto-detect spec type from draft GGUF metadata (#26814)
aic0d3r Aug 13, 2026
4a84b0a
metal : add TQ2_0 support (#26980)
ggerganov Aug 13, 2026
1d2869c
spec : auto-detect mtp draft model type (#27005)
CISC Aug 13, 2026
981184e
server : serve index.html with no-cache (#27006)
erusev Aug 13, 2026
2606220
chat : fix LFM2 tool call arg name prefix ambiguity (#26960)
aldehir Aug 13, 2026
a97123e
[SYCL] Support host pinned mem to improve SYCL Host-to-Device Memory …
arthw Aug 13, 2026
aee56b3
OpenVINO: Qwen3.5, memory optimization, and test-recurrent-state-roll…
wine99 Aug 13, 2026
9c5531e
ui: fix VITE_PUBLIC_SERVER variable reading (#24845)
brainrom Aug 13, 2026
fa4ec45
refactor: Naming (#27001)
allozaur Aug 13, 2026
bdffafa
ui: Refactor data-attrs constants, enum for bool strings (#27002)
ServeurpersoCom Aug 13, 2026
a94d563
common: apply CPU parameters across tools (#27026)
nikwen Aug 13, 2026
2bacf9e
dflash : clarify output logging of target_layer_ids (#27013)
danbev Aug 14, 2026
3d93885
sycl: fuse the gated-delta-net state writeback cpy (#26643)
Titaniumtown Aug 14, 2026
c6f6a92
ggml: force single thread on wasi (#25686)
MendyBerger Aug 14, 2026
6509138
sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779)
Titaniumtown Aug 14, 2026
885c5bb
tests : replace personal home directory paths with generic placeholde…
jimw567 Aug 14, 2026
77918ca
server: allow accessing /metrics and /slots during llama_decode() (#2…
ngxson Aug 14, 2026
4c1a0af
llama : allow virtual igpu devices (#26953)
ggerganov Aug 14, 2026
1692f9e
ggml : recurrent state rollback for ggml_ssm_scan (#26623)
lnigam Aug 14, 2026
06ae232
ggml : bump version to 0.20.0 (ggml/1584)
ggerganov Aug 14, 2026
9b05354
sync : ggml
ggerganov Aug 14, 2026
7e4c0a9
chat : pass reasoning_effort to template
sobakasu Aug 14, 2026
9e40df6
jinja : fix quadratic cost in gather_string_parts (#27034)
Yunzez Aug 14, 2026
6fed9f6
mtmd, common: various fixes (#27071)
ngxson Aug 14, 2026
16d222f
model : add support for MiniMaxText01ForCausalLM and MiniMaxM1ForCaus…
fairydreaming Aug 14, 2026
9d57ce4
mtmd: fix Granite4 Vision image sequence assembly (#26653)
hbattu73 Aug 14, 2026
7b38cb7
Fixed gating logic for problematic Intel driver version
rillomas Aug 14, 2026
6b4344e
fixed indent
rillomas Aug 14, 2026
0177dcc
common: migrate the deprecated --mmap/--no-mmap to --load-mode (#26934)
fboudra Aug 15, 2026
9b0a2ce
vulkan: add SHMEM_STRIDE_PAD/APPLY_SLM_A_RESHAPE for coopmat1 on Inte…
fish-jiang Aug 15, 2026
27df919
fix: check gguf array type before reading (#27075)
ngxson Aug 15, 2026
5f754ea
common: support --models-dir loading MTP assistant models (#24431)
EZForever Aug 15, 2026
77140d2
vendor : update cpp-httplib to 0.53.1 (#27103)
cabelo Aug 15, 2026
adb55e5
vendor: update BoringSSL to 0.20260813.0 (#27099)
cabelo Aug 15, 2026
22b8e31
server: re-design yield_to_queue thread model (#27133)
ngxson Aug 15, 2026
ad1de39
model: add Kimi-K3 text model (#26185)
pwilkin Aug 15, 2026
0d9ceae
ui: read structuredContent from MCP tool result when content is empty…
Gautam0507 Aug 15, 2026
ece963f
ui: mask API Key field in settings and error splash to stop browser a…
crowmoed Aug 15, 2026
10bf611
llama : check LoRA tensor data is within file bounds (#27056)
oakkaya Aug 16, 2026
b94041a
chat: refactor handling supports_string_content / supports_typed_cont…
ngxson Aug 16, 2026
3cb7ffb
model : remove some ggml_concat (#27176)
fairydreaming Aug 16, 2026
4df29be
ci : fix dry-run reporting in make-release job [no ci] (#27167)
danbev Aug 16, 2026
37a215c
[SYCL] support OP OPT_STEP_ADAMW, OPT_STEP_SGD (#25268)
arthw Aug 17, 2026
f275595
sycl: fix thread/block count in quantized cpy kernel launches (#27160)
Titaniumtown Aug 17, 2026
4695f00
llama-bench: fix deprecation warnings missing trailing newline (#27179)
fboudra Aug 17, 2026
cea66f4
ggml : bump version to 0.20.1 (ggml/1587)
ggerganov Aug 17, 2026
4197155
sync : ggml
ggerganov Aug 17, 2026
3733366
model : BailingMoE3 Support (#26608)
aetherbird Aug 17, 2026
fa88ae9
convert: add @ModelBase.example (#27208)
ngxson Aug 17, 2026
f9779dd
ci : make release workflows use a deploy key (#27229)
ggerganov Aug 17, 2026
7c35571
ci : allow make-release to target a specific commit (#27234)
ggerganov Aug 17, 2026
9f0d017
mtmd: harden preprocessor_granite (#27235)
ngxson Aug 17, 2026
d83f72d
ci : restore release.yml check during make-release.yml (#27247)
ggerganov Aug 17, 2026
7077abb
ui: add browser get_info tool (#27251)
ngxson Aug 17, 2026
9cd719a
model: support speculators-format checkpoints for DSpark (#26275)
wjinxu Aug 17, 2026
805984d
ci : reduce builds in build-xcframework.sh (#27252)
ggerganov Aug 17, 2026
666f889
ui: move get_datetime tool to frontend (#27255)
ngxson Aug 17, 2026
34af94c
ci : push release tag explicitly in release.yml (#27261)
ggerganov Aug 17, 2026
39be55c
vendor: move hash to vendor (#27262)
ngxson Aug 17, 2026
60eeeb6
cuda : skip UMA override for HIP builds (#27083)
superm1 Aug 17, 2026
b75ecd1
mtmd : skip thumbnail for non-tiled LFM2 images (#27246)
tdakhran Aug 17, 2026
d8df12e
vocab : support integer tokenizer scores (#27260)
CISC Aug 17, 2026
ed1c3a2
mtmd: use sha256 for input hashing (#27274)
ngxson Aug 17, 2026
533b182
server: save processed mtmd chunks as placeholder (#27278)
ngxson Aug 17, 2026
087f94d
doc: document MCP stdio servers and CORS defaults in the server READM…
ServeurpersoCom Aug 17, 2026
058df67
ci: more optimizations (#26983)
netrunnereve Aug 17, 2026
0021a77
ui: Refactor Built-In Tools naming (Server/Browser) (#27271)
allozaur Aug 17, 2026
01818e4
ui: enforce alphabetical enum member ordering (#27272)
ngxson Aug 17, 2026
25ae3a9
CUDA: MMVQ nwarps=8 for bs=1 for dense models on DGX Spark (#26843)
ynankani Aug 18, 2026
8b86400
ci : create pre-release with change log and nightly link in make-rele…
ggerganov Aug 18, 2026
27e345b
build : fix xcframework + cmake clean-up (#27304)
ggerganov Aug 18, 2026
da786dc
ggml : bump version to 0.20.2 (ggml/1589)
ggerganov Aug 18, 2026
1511ce3
sync : ggml
ggerganov Aug 18, 2026
7acdbb1
mtmd: fix LFM2 image tiling threshold (#27057)
BlackFoil Aug 18, 2026
c029602
ci: add Windows ARM64 CUDA support to the manual workflow (#27300)
shivamkumard-ctrl Aug 18, 2026
9d77fa1
ci : Update OpenVINO to 2026.3, skip nemotron-h rollback test (#27292)
wine99 Aug 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
20 changes: 10 additions & 10 deletions .devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
ARG OPENVINO_VERSION_MAJOR=2026.2.1
ARG OPENVINO_VERSION_FULL=2026.2.1.21919.ede283a88e3
ARG OPENVINO_VERSION_MAJOR=2026.3
ARG OPENVINO_VERSION_FULL=2026.3.0.22451.bd8d6542e3c
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.36.3
ARG IGC_VERSION_FULL=2_2.36.3+21719
ARG COMPUTE_RUNTIME_VERSION=26.22.38646.4
ARG COMPUTE_RUNTIME_VERSION_FULL=26.22.38646.4-0
ARG IGC_VERSION=v2.38.2
ARG IGC_VERSION_FULL=2_2.38.2+22051
ARG COMPUTE_RUNTIME_VERSION=26.27.39122.11
ARG COMPUTE_RUNTIME_VERSION_FULL=26.27.39122.11-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
ARG NPU_DRIVER_VERSION=v1.33.0
ARG NPU_DRIVER_FULL=v1.33.0.20260529-26625960453
ARG LIBZE1_VERSION=1.27.0-1~24.04~ppa2
ARG NPU_DRIVER_VERSION=v1.35.0
ARG NPU_DRIVER_FULL=v1.35.0.20260722-29947505341
ARG LIBZE1_VERSION=1.28.2-1~24.04~ppa1

# Optional proxy build arguments
ARG http_proxy=
Expand Down Expand Up @@ -170,7 +170,7 @@ RUN --mount=type=cache,target=/var/cache/intel-npu,sharing=locked \
fi; \
DEB=/var/cache/intel-npu/libze1_${LIBZE1_VERSION}_amd64.deb; \
if [ ! -f "$DEB" ]; then \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260324T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260606T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
fi; \
mkdir /tmp/npu/ && cd /tmp/npu/ && tar -xf "$TGZ" && cp "$DEB" .; \
apt-get update; \
Expand Down
20 changes: 0 additions & 20 deletions .github/actions/linux-setup-vulkan/action.yml

This file was deleted.

3 changes: 1 addition & 2 deletions .github/actions/windows-setup-cuda/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,7 @@ inputs:
required: true
cuda_arch:
description: "CUDA target architecture"
required: false
default: "x64"
required: true

runs:
using: "composite"
Expand Down
28 changes: 23 additions & 5 deletions .github/actions/windows-setup-rocm/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,26 @@ inputs:
runs:
using: "composite"
steps:
- name: Setup ROCm
uses: ./.github/actions/install-exe
with:
url: https://download.amd.com/developer/eula/rocm-hub/AMD-Software-PRO-Edition-${{ inputs.version }}-Win11-For-HIP.exe
args: -install
- name: Install ROCm with Wheels
shell: pwsh
run: |
$ErrorActionPreference = "Stop"
write-host "Setting up Python virtual environment"
# Create the venv directly at the cache location to avoid relocation issues
New-Item -Path "C:\TheRock\build" -ItemType Directory -Force | Out-Null
python -m venv C:\TheRock\build\.venv
& C:\TheRock\build\.venv\Scripts\Activate.ps1
write-host "Upgrading pip"
python -m pip install --upgrade pip
write-host "Installing ROCm wheels for multi-arch support"
# Install ROCm wheels for multi-arch support (this may take several minutes)
python -m pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "rocm[libraries,devel]==${{ inputs.version }}"
# Pre-expand the devel tree so it is included in the cache
write-host "Initializing ROCm devel tree"
rocm-sdk init
if ($LASTEXITCODE -ne 0) { throw "rocm-sdk init failed with exit code $LASTEXITCODE" }
write-host "Completed ROCm wheel installation to C:\TheRock\build"
85 changes: 29 additions & 56 deletions .github/workflows/build-cache.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,33 +10,6 @@ concurrency:
cancel-in-progress: true

jobs:
ubuntu-24-vulkan-cache:
runs-on: ubuntu-24.04

steps:
- name: Clone
id: checkout
uses: actions/checkout@v6

- name: Get latest Vulkan SDK version
id: vulkan_sdk_version
run: |
echo "VULKAN_SDK_VERSION=$(curl https://vulkan.lunarg.com/sdk/latest/linux.txt)" >> "$GITHUB_ENV"

- name: Setup Cache
uses: actions/cache@v5
id: cache-sdk
with:
path: ./vulkan_sdk
key: cache-gha-vulkan-sdk-${{ env.VULKAN_SDK_VERSION }}-${{ runner.os }}

- name: Setup Vulkan SDK
if: steps.cache-sdk.outputs.cache-hit != 'true'
uses: ./.github/actions/linux-setup-vulkan
with:
path: ./vulkan_sdk
version: ${{ env.VULKAN_SDK_VERSION }}

#ubuntu-24-spacemit-cache:
# runs-on: ubuntu-24.04

Expand Down Expand Up @@ -67,9 +40,9 @@ jobs:
runs-on: ubuntu-24.04

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"

steps:
- name: Clone
Expand All @@ -96,8 +69,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.2.1"
OPENVINO_VERSION_FULL: "2026.2.1.21919.ede283a88e3"
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"

steps:
- name: Clone
Expand All @@ -119,27 +92,27 @@ jobs:
version_major: ${{ env.OPENVINO_VERSION_MAJOR }}
version_full: ${{ env.OPENVINO_VERSION_FULL }}

windows-2022-rocm-cache:
runs-on: windows-2022

env:
# Make sure this is in sync with build.yml
HIPSDK_INSTALLER_VERSION: "26.Q1"

steps:
- name: Clone
id: checkout
uses: actions/checkout@v6

- name: Setup Cache
uses: actions/cache@v5
id: cache-rocm
with:
path: C:\Program Files\AMD\ROCm
key: cache-gha-rocm-${{ env.HIPSDK_INSTALLER_VERSION }}-${{ runner.os }}

- name: Setup ROCm
if: steps.cache-rocm.outputs.cache-hit != 'true'
uses: ./.github/actions/windows-setup-rocm
with:
version: ${{ env.HIPSDK_INSTALLER_VERSION }}
# windows-2022-rocm-cache:
# runs-on: windows-2022

# env:
# # Make sure this is in sync with release.yml and build-cuda-windows.yml
# ROCM_VERSION: "7.14.0"

# steps:
# - name: Clone
# id: checkout
# uses: actions/checkout@v6

# - name: Setup Cache
# uses: actions/cache@v5
# id: cache-rocm
# with:
# path: C:\TheRock\build
# key: rocm-wheels-${{ env.ROCM_VERSION }}-multi-arch-${{ runner.os }}

# - name: Setup ROCm
# if: steps.cache-rocm.outputs.cache-hit != 'true'
# uses: ./.github/actions/windows-setup-rocm
# with:
# version: ${{ env.ROCM_VERSION }}
14 changes: 10 additions & 4 deletions .github/workflows/build-cmake-pkg.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ on:

jobs:
linux:
runs-on: [self-hosted, Linux, CPU]
runs-on: [self-hosted, Linux]
steps:
- uses: actions/checkout@v6
with:
Expand All @@ -21,15 +21,21 @@ jobs:
-DLLAMA_BUILD_TOOLS=OFF \
-DLLAMA_BUILD_EXAMPLES=OFF \
-DLLAMA_BUILD_APP=OFF \
-DLLAMA_BUILD_IS_DEV=OFF \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release
cmake --build build --config Release -j $(nproc)
cmake --install build --prefix "$PREFIX" --config Release

export LLAMA_CONFIG="$PREFIX"/lib/cmake/llama/llama-config.cmake
tclsh <<'EOF'
set build(commit) [string trim [exec git rev-parse --short HEAD]]
set build(number) [string trim [exec git rev-list --count HEAD]]
set build(version) "0.0.$build(number)"

set cmakelists [read [open "CMakeLists.txt" r]]
regexp {set\(LLAMA_VERSION_MAJOR\s+(\d+)\)} $cmakelists -> major
regexp {set\(LLAMA_VERSION_MINOR\s+(\d+)\)} $cmakelists -> minor
regexp {set\(LLAMA_VERSION_PATCH\s+(\d+)\)} $cmakelists -> patch
set build(version) "$major.$minor.$patch"

set llamaconfig [read [open "$env(LLAMA_CONFIG)" r]]
set checks [list "set\\(LLAMA_VERSION \\s+$build(version)\\)" \
Expand All @@ -48,4 +54,4 @@ jobs:

cd examples/simple-cmake-pkg
cmake -S . -B build -DCMAKE_PREFIX_PATH="$PREFIX"/lib/cmake
cmake --build build
cmake --build build -j $(nproc)
18 changes: 4 additions & 14 deletions .github/workflows/build-cpu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ on:
paths: [
'.github/workflows/build-cpu.yml',
'.github/workflows/build-cmake-pkg.yml',
'ggml/src/ggml-rpc/**',
'**/CMakeLists.txt',
'**/.cmake',
'**/*.h',
Expand Down Expand Up @@ -94,8 +95,10 @@ jobs:
id: cmake_build
run: |
cmake -B build \
-DGGML_NATIVE=OFF \
-DLLAMA_FATAL_WARNINGS=ON \
-DGGML_RPC=ON
-DGGML_RPC=ON \
-DGGML_NATIVE=OFF
time cmake --build build --config Release -j $(nproc)

- name: Test
Expand All @@ -121,7 +124,6 @@ jobs:
env:
OPENBLAS_VERSION: 0.3.23
SDE_VERSION: 9.33.0-2024-01-07
VULKAN_VERSION: 1.4.357.0

strategy:
matrix:
Expand All @@ -132,9 +134,6 @@ jobs:
- build: 'x64-openblas'
arch: 'x64'
defines: '-G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_OPENMP=OFF -DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAS -DBLAS_INCLUDE_DIRS="$env:RUNNER_TEMP/openblas/include" -DBLAS_LIBRARIES="$env:RUNNER_TEMP/openblas/lib/openblas.lib"'
- build: 'x64-vulkan'
arch: 'x64'
defines: '-G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_VULKAN=ON'
- build: 'arm64'
arch: 'arm64'
defines: '-G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON'
Expand Down Expand Up @@ -165,15 +164,6 @@ jobs:
$lib = $(join-path $msvc 'bin\Hostx64\x64\lib.exe')
& $lib /machine:x64 "/def:${env:RUNNER_TEMP}/openblas/lib/libopenblas.def" "/out:${env:RUNNER_TEMP}/openblas/lib/openblas.lib" /name:openblas.dll

- name: Install Vulkan SDK
id: get_vulkan
if: ${{ matrix.build == 'x64-vulkan' }}
run: |
curl.exe -o $env:RUNNER_TEMP/VulkanSDK-Installer.exe -L "https://sdk.lunarg.com/sdk/download/${env:VULKAN_VERSION}/windows/vulkansdk-windows-X64-${env:VULKAN_VERSION}.exe"
& "$env:RUNNER_TEMP\VulkanSDK-Installer.exe" --accept-licenses --default-answer --confirm-command install
Add-Content $env:GITHUB_ENV "VULKAN_SDK=C:\VulkanSDK\${env:VULKAN_VERSION}"
Add-Content $env:GITHUB_PATH "C:\VulkanSDK\${env:VULKAN_VERSION}\bin"

- name: Install Ninja
id: install_ninja
run: |
Expand Down
Loading