Skip to content

Resolve llama-server before spawning; stop faking a GPU fallback - #183

Open
Endle wants to merge 1 commit into
masterfrom
fix/llama-server-resolution
Open

Endle wants to merge 1 commit into
masterfrom
fix/llama-server-resolution

Conversation

@Endle

@Endle Endle commented Aug 22, 2026

Copy link
Copy Markdown
Owner

The symptom

Starting the server with a working llama-server on PATH:

[WARN ] chat backend failed to start with -ngl 99 (failed to spawn llama-server:
        spawn chat backend: No such file or directory (os error 2)); falling back to CPU (-ngl 0)
[INFO ] spawning chat backend on port 8081: "llama-server" ... "-ngl" "0"
[ERROR] LLM backend failed to start: failed to spawn llama-server: spawn chat backend:
        No such file or directory (os error 2)

Two separate bugs are visible there.

1. We didn't resolve the binary the way the shell does

The PATH entry naming the binary was the literal four-character string ~/apps/bin, from a profile line that quotes the tilde:

export PATH="~/apps/bin:$PATH"   # tilde stays literal

Bash tilde-expands PATH elements itself during command lookup, so command -v llama-server resolves and running it at a prompt works. execvp(3) does no tilde expansion, so any process spawned from a program gets ENOENT — reproducible in one line:

python3 -c "subprocess.run(['llama-server','--version'])"
→ FileNotFoundError: [Errno 2] ... 'llama-server'

resolve_llama_server_bin now resolves --llama-server-bin up front, mirroring execvp's rule (a name with a separator is a path, anything else is a PATH search) but tilde-expanding each entry the way the shell does, and hands Command::new an absolute path. Skipped for .llamafile models, which are their own server.

2. Every failure was treated as a GPU failure

LlmError::Spawn conflated "could not start the process" (ENOENT, EACCES, uncreatable log file) with "child started, then died before serving /health". Only the second is the GPU-init signature the CPU fallback exists for. The first reproduces identically at -ngl 0, so the loop printed a misleading falling back to CPU and then failed with the same error — pointing at -ngl instead of at PATH.

Split into LlmError::Launch and LlmError::EarlyExit, and the fallback now retries only on EarlyExit / HealthCheck.

A missing binary reports itself once, with the fix in the message:

[ERROR] LLM backend failed to start: failed to launch llama-server:
        llama-server-that-does-not-exist not found in PATH. Install llama.cpp's
        llama-server and put it on PATH, pass --llama-server-bin /path/to/llama-server,
        or point --chat-endpoint at an already-running server. PATH=...

Verification

  • cargo test --all-targets — 98 passing, incl. 7 new unit tests covering the retry gate, PATH search (first match wins, non-executables skipped), literal-tilde expansion, and explicit-path resolution.
  • cargo clippy --locked — no new warnings.
  • End-to-end against the shell profile that triggered this, unmodified: spawning chat backend ... "/home/lizhenbo/apps/bin/llama-server" ... → chat backend ready at http://127.0.0.1:8081 (-ngl 99), first attempt.
  • End-to-end with a deliberately missing binary: single error, no bogus CPU fallback.

🤖 Generated with Claude Code

Two bugs, one symptom. Starting the server with a perfectly good
llama-server on PATH died with:

    [WARN ] chat backend failed to start with -ngl 99 (failed to spawn
            llama-server: ... No such file or directory (os error 2));
            falling back to CPU (-ngl 0)
    [ERROR] LLM backend failed to start: ... No such file or directory

The binary was there. The PATH entry naming it was the literal string
`~/apps/bin`, from a profile line that quotes the tilde
(`export PATH="~/bin:$PATH"`). Bash expands PATH elements itself during
command lookup, so `command -v llama-server` and typing it at a prompt
both work; `execvp` does no tilde expansion, so every spawn from a
program gets ENOENT. Resolve `--llama-server-bin` ourselves before the
first attempt — tilde-expanding each PATH entry the way the shell does —
and hand `Command::new` an absolute path.

The second bug is the warning above it. The retry loop treated *every*
failure as a GPU failure, so a missing binary produced a bogus "falling
back to CPU" and then failed again with the identical error, pointing at
-ngl instead of at PATH. `LlmError::Spawn` conflated "could not start the
process" with "child died before serving /health" — only the latter is
the GPU-init signature worth retrying. Split it into `Launch` and
`EarlyExit` and gate the fallback on the failure being GPU-shaped.

A missing binary now reports itself, once, with the fix in the message:

    [ERROR] LLM backend failed to start: failed to launch llama-server:
            llama-server-that-does-not-exist not found in PATH. Install
            llama.cpp's llama-server and put it on PATH, pass
            --llama-server-bin /path/to/llama-server, or point
            --chat-endpoint at an already-running server. PATH=...

Verified end-to-end against the unmodified shell profile that triggered
this: chat backend now comes up at -ngl 99 on the first attempt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant