Repository navigation
Conversation
Two bugs, one symptom. Starting the server with a perfectly good
llama-server on PATH died with:
[WARN ] chat backend failed to start with -ngl 99 (failed to spawn
llama-server: ... No such file or directory (os error 2));
falling back to CPU (-ngl 0)
[ERROR] LLM backend failed to start: ... No such file or directory
The binary was there. The PATH entry naming it was the literal string
`~/apps/bin`, from a profile line that quotes the tilde
(`export PATH="~/bin:$PATH"`). Bash expands PATH elements itself during
command lookup, so `command -v llama-server` and typing it at a prompt
both work; `execvp` does no tilde expansion, so every spawn from a
program gets ENOENT. Resolve `--llama-server-bin` ourselves before the
first attempt — tilde-expanding each PATH entry the way the shell does —
and hand `Command::new` an absolute path.
The second bug is the warning above it. The retry loop treated *every*
failure as a GPU failure, so a missing binary produced a bogus "falling
back to CPU" and then failed again with the identical error, pointing at
-ngl instead of at PATH. `LlmError::Spawn` conflated "could not start the
process" with "child died before serving /health" — only the latter is
the GPU-init signature worth retrying. Split it into `Launch` and
`EarlyExit` and gate the fallback on the failure being GPU-shaped.
A missing binary now reports itself, once, with the fix in the message:
[ERROR] LLM backend failed to start: failed to launch llama-server:
llama-server-that-does-not-exist not found in PATH. Install
llama.cpp's llama-server and put it on PATH, pass
--llama-server-bin /path/to/llama-server, or point
--chat-endpoint at an already-running server. PATH=...
Verified end-to-end against the unmodified shell profile that triggered
this: chat backend now comes up at -ngl 99 on the first attempt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The symptom
Starting the server with a working
llama-serveronPATH:Two separate bugs are visible there.
1. We didn't resolve the binary the way the shell does
The
PATHentry naming the binary was the literal four-character string~/apps/bin, from a profile line that quotes the tilde:Bash tilde-expands
PATHelements itself during command lookup, socommand -v llama-serverresolves and running it at a prompt works.execvp(3)does no tilde expansion, so any process spawned from a program gets ENOENT — reproducible in one line:resolve_llama_server_binnow resolves--llama-server-binup front, mirroringexecvp's rule (a name with a separator is a path, anything else is aPATHsearch) but tilde-expanding each entry the way the shell does, and handsCommand::newan absolute path. Skipped for.llamafilemodels, which are their own server.2. Every failure was treated as a GPU failure
LlmError::Spawnconflated "could not start the process" (ENOENT, EACCES, uncreatable log file) with "child started, then died before serving/health". Only the second is the GPU-init signature the CPU fallback exists for. The first reproduces identically at-ngl 0, so the loop printed a misleadingfalling back to CPUand then failed with the same error — pointing at-nglinstead of atPATH.Split into
LlmError::LaunchandLlmError::EarlyExit, and the fallback now retries only onEarlyExit/HealthCheck.A missing binary reports itself once, with the fix in the message:
Verification
cargo test --all-targets— 98 passing, incl. 7 new unit tests covering the retry gate,PATHsearch (first match wins, non-executables skipped), literal-tilde expansion, and explicit-path resolution.cargo clippy --locked— no new warnings.spawning chat backend ... "/home/lizhenbo/apps/bin/llama-server" ...→chat backend ready at http://127.0.0.1:8081 (-ngl 99), first attempt.🤖 Generated with Claude Code