ggml-org#27180
Summary
The frontend compiled-model cache introduced in ggml-org#26952 (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR) exports a CompiledModel whose Parameter/Result friendly names embed the #N suffix produced by get_tensor_ov_name(). That suffix is ggml_hash_find(&cgraph->visited_hash_set, tensor) — a slot index derived from the tensor's pointer value (ggml_hash is (uintptr_t)p >> 4). It therefore changes from process to process. On import in a fresh process, the decoder recomputes names from the new cgraph, gets different suffixes, and the very first input binding lookup throws std::out_of_range ("map::at"). The cache HIT loads fast but every inference fails. This is deterministic and affects any model whose graph has suffixed port names (KV-cache / recurrent-state inputs, i.e. every stateful LLM graph).
Environment
llama.cpp master, b10428 (885c5bb), GGML_OPENVINO=ON, BUILD_SHARED_LIBS=OFF
OpenVINO 2026.4.0 (openvino-git r263.g0f453eb8dca), GPU plugin, intel-compute-runtime 26.27.39122.11
Intel Core Ultra 7 356H (Panther Lake) iGPU, Arch Linux
Model: qwen35-arch GGUF (GatedDeltaNet linear attention, 9B Q4_K_M), -ngl 99, -c 16384
Repro
# 1. fresh cache dir: first start compiles + exports the blob
env GGML_OPENVINO_DEVICE=GPU \
GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/tmp/ov-test \
llama-server -m model.gguf --port 18100 -c 16384 -ngl 99 -dev OPENVINO0
# log: "ggml-openvino: model cache WROTE /tmp/ov-test/<fp>.blob"
# /completion works. Stop the server.
# 2. second start: cache HIT, fast load
env GGML_OPENVINO_DEVICE=GPU \
GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/tmp/ov-test \
llama-server -m model.gguf --port 18100 -c 16384 -ngl 99 -dev OPENVINO0
# log: "ggml-openvino: model cache HIT /tmp/ov-test/<fp>.blob"
curl localhost:18100/completion -d '{"prompt":"1+1=","n_predict":8,"temperature":0}'
# => HTTP 500 "Compute error", every time
Note: the backend's cache lines are GGML_LOG_INFO and are suppressed at the
server's default log verbosity — pass -lv 5 (as in the invocations above) to
see the WROTE/HIT lines.
Server log on the HIT path:
I ggml-openvino: model cache HIT /tmp/ov-test/<fp>.blob
E GGML OpenVINO backend std::exception: map::at
E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
E srv decode: Compute error. off = 0, n_batch = 2048, ret = -3
Root cause (gdb-confirmed)
Backtrace of the throw on the import path (first inference after HIT,
RelWithDebInfo build, catch throw on std::out_of_range):
std::__throw_out_of_range
std::map<std::string, ggml_tensor*>::at(__k="cache_k_l11#55152") stl_map.h:611
GgmlOvDecoder::get_input_ggml_tensor(name="cache_k_l11#55152") ggml-decoder.h:218 <- m_inputs.at(name)
convert_ggml_input_to_ov utils.cpp:1028
get_ov_input_tensor utils.cpp:1069
ov_graph_compute_dynamic (binding loop over cm.inputs()) utils.cpp:506
The imported model's input port friendly name is cache_k_l11#55152. The
suffix is appended by get_tensor_ov_name() (ggml-decoder.cpp):
const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
if (((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || is_kvcache(tensor, nullptr)) && ...) {
return std::string(tensor->name) + "#" + std::to_string(hash_pos);
}
hash_pos is a slot in a pointer-keyed hash set; across two runs of the same
binary + model we observed cache_k_l11#2940 and cache_k_l11#55152. The
compile path works because both the ov::Model and the decoder are built from
the same cgraph in the same process. The import path compares names baked at
export time against names recomputed in a different process, so m_inputs
has no such key and .at() throws.
This means the feature could not have worked across processes for any graph
with KV-cache/recurrent-state inputs — only same-process round-trips or graphs
without suffixed port names would survive.
Proposed fix
Make the disambiguation suffix deterministic across processes: use the
tensor's index in the cgraph nodes/leafs arrays (build order is
deterministic for a given model + graph params) instead of the pointer-hash
slot:
static std::string get_tensor_ov_name(const ggml_cgraph * cgraph, const ggml_tensor * tensor) {
if (tensor == nullptr) {
return "";
}
if ((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) {
const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
if (hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
for (int i = 0; i < cgraph->n_nodes; i++) {
if (cgraph->nodes[i] == tensor) {
return std::string(tensor->name) + "#" + std::to_string(i);
}
}
for (int i = 0; i < cgraph->n_leafs; i++) {
if (cgraph->leafs[i] == tensor) {
return std::string(tensor->name) + "#l" + std::to_string(i);
}
}
}
}
return tensor->name;
}
Two supporting changes:
- Cache-format version in the fingerprint.
ggml_openvino_model_cache_extra_cfg() should fold in a naming-scheme version so blobs exported by older builds MISS instead of importing with stale names (otherwise: HIT + map::at, the poisoned-blob trap).
GgmlOvDecoder::collect_weight_names() should mirror create_weight_nodes() exactly — it is missing the is_mul_mat_id_expert_weight(node, i) term, so on the import path non-quantized MoE expert weights are misclassified as model inputs (harmless for binding today, but the two selectors must agree by construction).
Working patch (verified locally: compile path WROTE + inference OK; import
path HIT + inference OK) is attached below. Verification detail: with the
patch, the 9B model compiles in ~30 s and imports in ~17 s; a fresh-process
import returns the same completion output as the compile path, repeated
across several processes, with no map::at / Compute error.
Notes:
- The O(n) scan per call is bounded by decoder construction (per unique graph shape per process), not per decode step; measured no load-time regression (9B model: import ~17 s before/after).
- The lookup that tolerates misses on the output side (get_model_outputs().find(...)) silently leaves misnamed outputs unbound — with the old naming this would have produced wrong results instead of an error, had the input side not thrown first.
- Unrelated observation while debugging: GGML_OPENVINO_DEBUG_INPUT=1 segfaults during binding for this model on both compile and import paths (separate manifest; happy to file separately with a backtrace).
Patch
diff --git a/ggml/src/ggml-openvino/ggml-decoder.cpp b/ggml/src/ggml-openvino/ggml-decoder.cpp
index 599f41aeb..8a3f64ee8 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.cpp
+++ b/ggml/src/ggml-openvino/ggml-decoder.cpp
@@ -180,10 +180,31 @@ static std::string get_tensor_ov_name(const ggml_cgraph * cgraph, const ggml_ten
if (tensor == nullptr) {
return "";
}
- const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
- if (((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) &&
- hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
- return std::string(tensor->name) + "#" + std::to_string(hash_pos);
+ if ((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) {
+ const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
+ if (hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
+ // Disambiguate same-named tensors with a deterministic, process-independent
+ // ordinal: the tensor's index in the cgraph nodes/leafs arrays. The graph
+ // build order is fixed for a given model + graph params, so these names are
+ // stable across processes. That stability is required by the frontend
+ // compiled-model cache (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR): the exported
+ // model's Parameter/Result friendly names are baked in at export time and
+ // must match the names recomputed from a fresh cgraph in the importing
+ // process. The previous hash_pos suffix derives from the tensor's pointer
+ // value (ggml_hash) and changes from process to process, so an imported
+ // model's port names (e.g. "cache_k_l11#55152") never matched the decoder's
+ // input map and binding failed with std::out_of_range ("map::at").
+ for (int i = 0; i < cgraph->n_nodes; i++) {
+ if (cgraph->nodes[i] == tensor) {
+ return std::string(tensor->name) + "#" + std::to_string(i);
+ }
+ }
+ for (int i = 0; i < cgraph->n_leafs; i++) {
+ if (cgraph->leafs[i] == tensor) {
+ return std::string(tensor->name) + "#l" + std::to_string(i);
+ }
+ }
+ }
}
return tensor->name;
}
@@ -1045,7 +1066,8 @@ std::set<std::string> GgmlOvDecoder::collect_weight_names(ggml_cgraph * cgraph)
}
if (!src->view_src) {
ggml_backend_buffer * buffer = src->buffer;
- if (buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS || ggml_is_quantized(src->type)) {
+ if (buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS || ggml_is_quantized(src->type) ||
+ is_mul_mat_id_expert_weight(node, i)) {
names.insert(src_name);
}
}
diff --git a/ggml/src/ggml-openvino/utils.cpp b/ggml/src/ggml-openvino/utils.cpp
index 4df8381dc..50266c781 100644
--- a/ggml/src/ggml-openvino/utils.cpp
+++ b/ggml/src/ggml-openvino/utils.cpp
@@ -142,6 +142,11 @@ static uint64_t ggml_openvino_model_cache_extra_cfg(const std::string & device,
device == "GPU";
uint64_t extra_cfg = 0;
+ // Cache-format version, folded into the fingerprint so blobs exported by
+ // builds with incompatible port naming MISS instead of importing stale names.
+ // v2: deterministic cgraph-ordinal tensor names in get_tensor_ov_name()
+ // (previously a pointer-derived hash_pos suffix that broke import).
+ extra_cfg = extra_cfg * 131 + 2u;
extra_cfg = extra_cfg * 131 + (stateful ? 1u : 0u);
extra_cfg = extra_cfg * 131 + (ggml_openvino_reduce_compile_mem_enabled() ? 1u : 0u);
extra_cfg = extra_cfg * 131 + (ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_KV_SLICE") ? 1u : 0u);
ggml-org#27180
Summary
The frontend compiled-model cache introduced in ggml-org#26952 (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR) exports a CompiledModel whose Parameter/Result friendly names embed the #N suffix produced by get_tensor_ov_name(). That suffix is ggml_hash_find(&cgraph->visited_hash_set, tensor) — a slot index derived from the tensor's pointer value (ggml_hash is (uintptr_t)p >> 4). It therefore changes from process to process. On import in a fresh process, the decoder recomputes names from the new cgraph, gets different suffixes, and the very first input binding lookup throws std::out_of_range ("map::at"). The cache HIT loads fast but every inference fails. This is deterministic and affects any model whose graph has suffixed port names (KV-cache / recurrent-state inputs, i.e. every stateful LLM graph).
Environment
llama.cpp master, b10428 (885c5bb), GGML_OPENVINO=ON, BUILD_SHARED_LIBS=OFF
OpenVINO 2026.4.0 (openvino-git r263.g0f453eb8dca), GPU plugin, intel-compute-runtime 26.27.39122.11
Intel Core Ultra 7 356H (Panther Lake) iGPU, Arch Linux
Model: qwen35-arch GGUF (GatedDeltaNet linear attention, 9B Q4_K_M), -ngl 99, -c 16384
Repro
Note: the backend's cache lines are GGML_LOG_INFO and are suppressed at the
server's default log verbosity — pass -lv 5 (as in the invocations above) to
see the WROTE/HIT lines.
Server log on the HIT path:
Root cause (gdb-confirmed)
Backtrace of the throw on the import path (first inference after HIT,
RelWithDebInfo build,
catch throwonstd::out_of_range):The imported model's input port friendly name is
cache_k_l11#55152. Thesuffix is appended by
get_tensor_ov_name()(ggml-decoder.cpp):hash_posis a slot in a pointer-keyed hash set; across two runs of the samebinary + model we observed
cache_k_l11#2940andcache_k_l11#55152. Thecompile path works because both the ov::Model and the decoder are built from
the same cgraph in the same process. The import path compares names baked at
export time against names recomputed in a different process, so
m_inputshas no such key and
.at()throws.This means the feature could not have worked across processes for any graph
with KV-cache/recurrent-state inputs — only same-process round-trips or graphs
without suffixed port names would survive.
Proposed fix
Make the disambiguation suffix deterministic across processes: use the
tensor's index in the cgraph nodes/leafs arrays (build order is
deterministic for a given model + graph params) instead of the pointer-hash
slot:
Two supporting changes:
ggml_openvino_model_cache_extra_cfg()should fold in a naming-scheme version so blobs exported by older builds MISS instead of importing with stale names (otherwise: HIT +map::at, the poisoned-blob trap).GgmlOvDecoder::collect_weight_names()should mirrorcreate_weight_nodes()exactly — it is missing the is_mul_mat_id_expert_weight(node, i) term, so on the import path non-quantized MoE expert weights are misclassified as model inputs (harmless for binding today, but the two selectors must agree by construction).Working patch (verified locally: compile path WROTE + inference OK; import
path HIT + inference OK) is attached below. Verification detail: with the
patch, the 9B model compiles in ~30 s and imports in ~17 s; a fresh-process
import returns the same completion output as the compile path, repeated
across several processes, with no map::at / Compute error.
Notes:
Patch