Skip to content

[ICS] frontend compiled-model cache import loads but cannot compute. #298

Description

@haarika-madaka

ggml-org#27180

Summary
The frontend compiled-model cache introduced in ggml-org#26952 (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR) exports a CompiledModel whose Parameter/Result friendly names embed the #N suffix produced by get_tensor_ov_name(). That suffix is ggml_hash_find(&cgraph->visited_hash_set, tensor) — a slot index derived from the tensor's pointer value (ggml_hash is (uintptr_t)p >> 4). It therefore changes from process to process. On import in a fresh process, the decoder recomputes names from the new cgraph, gets different suffixes, and the very first input binding lookup throws std::out_of_range ("map::at"). The cache HIT loads fast but every inference fails. This is deterministic and affects any model whose graph has suffixed port names (KV-cache / recurrent-state inputs, i.e. every stateful LLM graph).

Environment
llama.cpp master, b10428 (885c5bb), GGML_OPENVINO=ON, BUILD_SHARED_LIBS=OFF
OpenVINO 2026.4.0 (openvino-git r263.g0f453eb8dca), GPU plugin, intel-compute-runtime 26.27.39122.11
Intel Core Ultra 7 356H (Panther Lake) iGPU, Arch Linux
Model: qwen35-arch GGUF (GatedDeltaNet linear attention, 9B Q4_K_M), -ngl 99, -c 16384

Repro

# 1. fresh cache dir: first start compiles + exports the blob
env GGML_OPENVINO_DEVICE=GPU \
    GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/tmp/ov-test \
    llama-server -m model.gguf --port 18100 -c 16384 -ngl 99 -dev OPENVINO0
# log: "ggml-openvino: model cache WROTE /tmp/ov-test/<fp>.blob"
# /completion works. Stop the server.

# 2. second start: cache HIT, fast load
env GGML_OPENVINO_DEVICE=GPU \
    GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/tmp/ov-test \
    llama-server -m model.gguf --port 18100 -c 16384 -ngl 99 -dev OPENVINO0
# log: "ggml-openvino: model cache HIT /tmp/ov-test/<fp>.blob"
curl localhost:18100/completion -d '{"prompt":"1+1=","n_predict":8,"temperature":0}'
# => HTTP 500 "Compute error", every time

Note: the backend's cache lines are GGML_LOG_INFO and are suppressed at the
server's default log verbosity — pass -lv 5 (as in the invocations above) to
see the WROTE/HIT lines.

Server log on the HIT path:

I ggml-openvino: model cache HIT /tmp/ov-test/<fp>.blob
E GGML OpenVINO backend std::exception: map::at
E graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
E srv  decode: Compute error. off = 0, n_batch = 2048, ret = -3

Root cause (gdb-confirmed)
Backtrace of the throw on the import path (first inference after HIT,
RelWithDebInfo build, catch throw on std::out_of_range):

std::__throw_out_of_range
std::map<std::string, ggml_tensor*>::at(__k="cache_k_l11#55152")   stl_map.h:611
GgmlOvDecoder::get_input_ggml_tensor(name="cache_k_l11#55152")     ggml-decoder.h:218   <- m_inputs.at(name)
convert_ggml_input_to_ov                                           utils.cpp:1028
get_ov_input_tensor                                                utils.cpp:1069
ov_graph_compute_dynamic  (binding loop over cm.inputs())          utils.cpp:506

The imported model's input port friendly name is cache_k_l11#55152. The
suffix is appended by get_tensor_ov_name() (ggml-decoder.cpp):

const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
if (((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || is_kvcache(tensor, nullptr)) && ...) {
    return std::string(tensor->name) + "#" + std::to_string(hash_pos);
}

hash_pos is a slot in a pointer-keyed hash set; across two runs of the same
binary + model we observed cache_k_l11#2940 and cache_k_l11#55152. The
compile path works because both the ov::Model and the decoder are built from
the same cgraph in the same process. The import path compares names baked at
export time against names recomputed in a different process, so m_inputs
has no such key and .at() throws.

This means the feature could not have worked across processes for any graph
with KV-cache/recurrent-state inputs — only same-process round-trips or graphs
without suffixed port names would survive.

Proposed fix

Make the disambiguation suffix deterministic across processes: use the
tensor's index in the cgraph nodes/leafs arrays (build order is
deterministic for a given model + graph params) instead of the pointer-hash
slot:

static std::string get_tensor_ov_name(const ggml_cgraph * cgraph, const ggml_tensor * tensor) {
    if (tensor == nullptr) {
        return "";
    }
    if ((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) {
        const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
        if (hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
            for (int i = 0; i < cgraph->n_nodes; i++) {
                if (cgraph->nodes[i] == tensor) {
                    return std::string(tensor->name) + "#" + std::to_string(i);
                }
            }
            for (int i = 0; i < cgraph->n_leafs; i++) {
                if (cgraph->leafs[i] == tensor) {
                    return std::string(tensor->name) + "#l" + std::to_string(i);
                }
            }
        }
    }
    return tensor->name;
}

Two supporting changes:

  1. Cache-format version in the fingerprint. ggml_openvino_model_cache_extra_cfg() should fold in a naming-scheme version so blobs exported by older builds MISS instead of importing with stale names (otherwise: HIT + map::at, the poisoned-blob trap).
  2. GgmlOvDecoder::collect_weight_names() should mirror create_weight_nodes() exactly — it is missing the is_mul_mat_id_expert_weight(node, i) term, so on the import path non-quantized MoE expert weights are misclassified as model inputs (harmless for binding today, but the two selectors must agree by construction).

Working patch (verified locally: compile path WROTE + inference OK; import
path HIT + inference OK) is attached below. Verification detail: with the
patch, the 9B model compiles in ~30 s and imports in ~17 s; a fresh-process
import returns the same completion output as the compile path, repeated
across several processes, with no map::at / Compute error.

Notes:

  • The O(n) scan per call is bounded by decoder construction (per unique graph shape per process), not per decode step; measured no load-time regression (9B model: import ~17 s before/after).
  • The lookup that tolerates misses on the output side (get_model_outputs().find(...)) silently leaves misnamed outputs unbound — with the old naming this would have produced wrong results instead of an error, had the input side not thrown first.
  • Unrelated observation while debugging: GGML_OPENVINO_DEBUG_INPUT=1 segfaults during binding for this model on both compile and import paths (separate manifest; happy to file separately with a backtrace).

Patch

diff --git a/ggml/src/ggml-openvino/ggml-decoder.cpp b/ggml/src/ggml-openvino/ggml-decoder.cpp
index 599f41aeb..8a3f64ee8 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.cpp
+++ b/ggml/src/ggml-openvino/ggml-decoder.cpp
@@ -180,10 +180,31 @@ static std::string get_tensor_ov_name(const ggml_cgraph * cgraph, const ggml_ten
     if (tensor == nullptr) {
         return "";
     }
-    const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
-    if (((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) &&
-        hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
-        return std::string(tensor->name) + "#" + std::to_string(hash_pos);
+    if ((tensor->flags & GGML_TENSOR_FLAG_COMPUTE) || GgmlOvDecoder::is_kvcache(tensor, nullptr)) {
+        const size_t hash_pos = ggml_hash_find(&cgraph->visited_hash_set, tensor);
+        if (hash_pos != GGML_HASHSET_FULL && ggml_bitset_get(cgraph->visited_hash_set.used, hash_pos)) {
+            // Disambiguate same-named tensors with a deterministic, process-independent
+            // ordinal: the tensor's index in the cgraph nodes/leafs arrays. The graph
+            // build order is fixed for a given model + graph params, so these names are
+            // stable across processes. That stability is required by the frontend
+            // compiled-model cache (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR): the exported
+            // model's Parameter/Result friendly names are baked in at export time and
+            // must match the names recomputed from a fresh cgraph in the importing
+            // process. The previous hash_pos suffix derives from the tensor's pointer
+            // value (ggml_hash) and changes from process to process, so an imported
+            // model's port names (e.g. "cache_k_l11#55152") never matched the decoder's
+            // input map and binding failed with std::out_of_range ("map::at").
+            for (int i = 0; i < cgraph->n_nodes; i++) {
+                if (cgraph->nodes[i] == tensor) {
+                    return std::string(tensor->name) + "#" + std::to_string(i);
+                }
+            }
+            for (int i = 0; i < cgraph->n_leafs; i++) {
+                if (cgraph->leafs[i] == tensor) {
+                    return std::string(tensor->name) + "#l" + std::to_string(i);
+                }
+            }
+        }
     }
     return tensor->name;
 }
@@ -1045,7 +1066,8 @@ std::set<std::string> GgmlOvDecoder::collect_weight_names(ggml_cgraph * cgraph)
             }
             if (!src->view_src) {
                 ggml_backend_buffer * buffer = src->buffer;
-                if (buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS || ggml_is_quantized(src->type)) {
+                if (buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS || ggml_is_quantized(src->type) ||
+                    is_mul_mat_id_expert_weight(node, i)) {
                     names.insert(src_name);
                 }
             }
diff --git a/ggml/src/ggml-openvino/utils.cpp b/ggml/src/ggml-openvino/utils.cpp
index 4df8381dc..50266c781 100644
--- a/ggml/src/ggml-openvino/utils.cpp
+++ b/ggml/src/ggml-openvino/utils.cpp
@@ -142,6 +142,11 @@ static uint64_t ggml_openvino_model_cache_extra_cfg(const std::string & device,
                                         device == "GPU";
 
     uint64_t extra_cfg = 0;
+    // Cache-format version, folded into the fingerprint so blobs exported by
+    // builds with incompatible port naming MISS instead of importing stale names.
+    // v2: deterministic cgraph-ordinal tensor names in get_tensor_ov_name()
+    //     (previously a pointer-derived hash_pos suffix that broke import).
+    extra_cfg = extra_cfg * 131 + 2u;
     extra_cfg = extra_cfg * 131 + (stateful ? 1u : 0u);
     extra_cfg = extra_cfg * 131 + (ggml_openvino_reduce_compile_mem_enabled() ? 1u : 0u);
     extra_cfg = extra_cfg * 131 + (ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_KV_SLICE") ? 1u : 0u);

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions