Repository navigation
Conversation
…erf pass)
## Security
- **Shepherd argv is `--`-terminated.** A command whose name starts with `-`
(`--cgroup-path`, `--kill-timeout`, `--pty`) could be parsed as a shepherd
flag and re-target the cgroup or stretch the kill ladder behind the API.
- **cgroup teardown requires `cgroup.kill`.** `cgroup_setup` probes the new
leaf with `access(W_OK)` and fails the spawn closed (rolling back its own
mkdir) when it is unavailable; without it a `setsid()` daemoniser escapes
`kill(-pgid)`. The cleanup write is now checked and logged.
- **Signal routing is probe-then-signal.** `CMD_KILL` over the UDS while the
shepherd Port lives; once the Port is dead or the send fails,
`NetRunner.Process` probes `kill(pid, 0)` before `nif_kill`. The Watcher's
port-liveness gate was removed: the Port is owned by the GenServer whose
death the Watcher handles, so it was always already closed (a no-op).
## Fixed
- `send_fds` restarted its 5 s `POLLOUT` window on every `EAGAIN`; it and
`write_fully` now share `wait_pollout()` with one lazily-computed absolute
deadline and a hard poll cap when the monotonic clock is unavailable.
- `kill/2` reports lost signals: `send_shepherd_command` returns
`:ok | {:error, _}`, failed sends fall back to the direct probe-and-kill,
and `{:error, :not_running}` replaces a misleading `:ok`.
`set_window_size/3` surfaces a failed send.
- Option validation is uniform: `Daemon.start_link/1` validates
`:process_opts` in the caller via the new public
`Process.validate_opts!/1`; `:kill_timeout` is range-checked (1..60_000)
and `:cgroup_path` rejects NUL bytes and non-binaries.
- `Protocol.kill/1` rejects signals outside 1..255 instead of truncating.
- macOS test flake: `sleep` markers built from `System.unique_integer/1`
crossed `INT32_MAX` late in a run; `sleep_marker/0` keeps them in range.
## Added
- `Stats.shepherd_error` exposes a post-spawn shepherd `MSG_ERROR`.
- Tests: Watcher kills an orphan after shepherd death, direct `kill/2` after
shepherd death, `Protocol` encoder byte layouts, post-spawn `MSG_ERROR`
recording, Daemon client-side validation, `:kill_timeout` validation,
`--` terminator. PTY-resize sentinel no longer matches its own failure
message; teardown test excludes the stdin forwarder from its drain sample.
- CI: `NR_CGROUP_DELEGATED=1` on the delegated Ubuntu leg makes the cgroup
tests' fail-closed branch a hard failure. Alpine leg documented as
fail-closed-only.
## Docs
- `NetRunner.Process.Nif` -> `NetRunner.Nif` and `Exec.parse_uds_message`
-> `Protocol.parse_uds_message` references corrected; Layer 3 is the NIF
owner monitor, not GC; pre-existing cgroup dir is a spawn error.
- `docs/backpressure.md`: `pipe-user-pages-soft` caveat (3 x 1 MiB per spawn
exhausts the default budget at ~21 concurrent commands; unmeasured).
- `bench/LINUX_VERIFICATION.md`: cgroup spawn-latency and >=32-concurrent
chunk-count rows for a Linux host to fill in.
## Deliberate deviations from triage
- Force-exit backstop kept at 5 s (not immediate): `recv_uds` can hit
`:ealready` with a real `MSG_CHILD_EXITED` still pending (PR #5 race).
- Leading-`-` commands are not rejected; the `--` terminator alone closes
the hole and `-foo` is a legal command name.
- No `{:error, {:shepherd_unreachable, _}}` shape; all send failures fall
through to the direct probe-and-kill.
Verified: format, `WERROR=1` compile (clang), credo --strict, 253 tests
green on two seeds. Not exercised: `:linux_only` cgroup tests, gcc build.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Second-pass fixes for the 2026-10-07 review of #12 (
1116a77), which added awhole-file security and performance audit of
c_src/. 14 findings fixed,3 triage items consciously narrowed (see bottom). Review artefacts live under
.claude/plans/review/(not in this PR).Security
---terminated. A command whose name starts with-(
--cgroup-path,--kill-timeout,--pty) could be parsed as a shepherdflag and re-target the cgroup or stretch the kill ladder behind the API.
cgroup.kill.cgroup_setupprobes the newleaf with
access(W_OK)and fails the spawn closed (rolling back its ownmkdir) when it is unavailable; without it asetsid()daemoniser escapeskill(-pgid). The cleanup write is now checked and logged.CMD_KILLover the UDS while theshepherd Port lives; once the Port is dead or the send fails,
NetRunner.Processprobeskill(pid, 0)beforenif_kill. The Watcher'sport-liveness gate was removed: the Port is owned by the GenServer whose
death the Watcher handles, so it was always already closed — a no-op that
documented a guarantee the code did not have.
Fixed
send_fdsrestarted its 5 sPOLLOUTwindow on everyEAGAIN; it andwrite_fullynow sharewait_pollout()with one lazily-computed absolutedeadline (a first-try write never reads the clock) and a hard poll cap when
the monotonic clock is unavailable.
kill/2reports lost signals:send_shepherd_commandreturns:ok | {:error, _}, failed sends fall back to the direct probe-and-kill,and
{:error, :not_running}replaces a misleading:ok.set_window_size/3surfaces a failed send.Daemon.start_link/1validates:process_optsin the caller via the new publicProcess.validate_opts!/1(previously raised inside
init/1and exited the linked caller);:kill_timeoutis range-checked (1..60_000) and:cgroup_pathrejectsNUL bytes and non-binaries — instead of a silent 10 s
:shepherd_connect_timeout.Protocol.kill/1rejects signals outside1..255instead of truncating.sleepmarkers built fromSystem.unique_integer/1crossed
INT32_MAXlate in a run and/bin/sleepexited 1;sleep_marker/0keeps them in range.Added
Stats.shepherd_errorexposes a post-spawn shepherdMSG_ERROR(was buriedin server state).
direct
kill/2after shepherd death,Protocolencoder byte layoutspinned to
protocol.h, post-spawnMSG_ERRORrecording, Daemonclient-side validation,
:kill_timeoutvalidation,--terminatorend-to-end. PTY-resize sentinel no longer matches its own failure message
(
"NEVER RESIZED" =~ "RESIZED"); teardown test excludes the stdinforwarder from its drain-task sample.
NR_CGROUP_DELEGATED=1on the delegated Ubuntu leg turns the cgrouptests' fail-closed branch into a hard failure, so a delegation regression
can't ship as a green build with zero cgroup coverage. Alpine leg
documented as fail-closed-only.
Docs
NetRunner.Process.Nif→NetRunner.NifandExec.parse_uds_message→Protocol.parse_uds_messagereferences corrected (CLAUDE.md,docs/modules.md,docs/protocol.md); Layer 3 is the NIF owner monitor,not GC; a pre-existing cgroup dir is a spawn error.
docs/backpressure.md:pipe-user-pages-softcaveat — 3 × 1 MiB perspawn exhausts the default budget at ~21 concurrent commands, after which
pipes fall back to 64 KiB. Derived from kernel defaults, unmeasured.
bench/LINUX_VERIFICATION.md: cgroup spawn-latency and ≥32-concurrentchunk-count rows for a Linux host to fill in.
Deliberate deviations from triage
recv_udscan return:ealreadywhile a genuineMSG_CHILD_EXITEDisstill pending as a
$socketmessage (the Retry UDS drain on port exit to fix macOS CI race #5 macOS race); synthesising 137immediately would overwrite a real exit code. Probe-then-signal already
removed the pid-reuse exposure from that window.
-commands are not rejected. The--terminator alone closesthe hole (BEAM-emitted flags can never equal
--), and-foois a legalcommand name.
{:error, {:shepherd_unreachable, _}}reply shape. All sendfailures fall through to the direct probe-and-kill; the observable error
is
{:error, :not_running}.Known follow-ups (second-pass review, not in this PR)
shepherd.c:913,1097) andnif_dup_fd(net_runner_nif.c:515)are still non-CLOEXEC — now flagged by convention C001.
Daemon.start_link(args: [])with no:cmdstill raisesKeyErrorinsideinit/1.kill -KILL <ppid>in two tests is unguarded on a reparented child;sleep_markercan prefix-collide underpgrep -f.Verification
mix format --check-formatted·WERROR=1 mix compile --warnings-as-errors(clang) ·
mix credo --strict— cleanmix test: 253 passed, 2 excluded (:linux_only), green on default seedand seed
151793(the one that previously failed)run(["--pty", "sh", "-c", "echo pwned"])→{"", 127, "execvp: --pty\n"}; kill via shepherd →:ok/ 137; kill afterexit →
{:error, :not_running}:linux_onlycgroup tests (incl. the newcgroup.killprobe), gcc-Werrorbuild, dialyzer