Skip to content

[Browser plugin] playwright run-server is never killed on Debian/Ubuntu (/bin/sh = dash): orphaned node process after every run, terminal hangs when output is piped #1754

Description

@likemusic

Description

After every pest run that includes browser tests, the playwright run-server node process
survives the pest process on Debian/Ubuntu-based environments (including Laravel Sail
runtimes). One orphan is leaked per run.

The most visible symptom: when pest output is piped (pest ... 2>&1 | cat, CI log
collectors, etc.), the terminal hangs forever after the final test summary — the orphaned
server inherits the stdout pipe, so the reader never gets EOF. Without a pipe the prompt
returns, but orphaned node processes silently accumulate until the machine/container is
restarted.

Environment

  • pest-plugin-browser v4.3.1 (the relevant code is unchanged in 4.x HEAD)
  • pest 4.x, PHP 8.3 (NTS), Laravel Sail runtime 8.3 (Ubuntu, /bin/sh → dash)
  • playwright 1.59.x, Chromium headless
  • Not reproducible on macOS (see root cause — that's why it is easy to miss upstream)

Root cause

ServerManager::playwright() builds the server command as a string and
PlaywrightNpmServer::start() runs it via Process::fromShellCommandline(). A string
command goes through proc_open → /bin/sh -c '...', and Symfony Process only prepends
exec for array commandlines — not for fromShellCommandline().

On Debian/Ubuntu /bin/sh is dash, which does not exec-replace itself with a single
-c command. The resulting process tree is:

php (pest)
 └─ sh -c ./node_modules/.bin/playwright run-server ...    ← Symfony Process child (getPid())
     └─ node ./node_modules/.bin/playwright run-server ... ← the actual server

PlaywrightNpmServer::stop() calls $process->stop(timeout: 0.1, signal: SIGTERM), which
signals only the direct child (sh). sh dies, the node server is reparented to PID 1
and keeps running forever. (The node process itself handles SIGTERM fine — killing it
directly works instantly.)

On macOS /bin/sh is bash, which does exec-replace a single -c command, so the signal
reaches node directly and everything shuts down cleanly — the bug is invisible there.

Standalone reproduction (no pest involved)

<?php
require __DIR__.'/vendor/autoload.php';

use Symfony\Component\Process\Process;

$process = Process::fromShellCommandline(
    './node_modules/.bin/playwright run-server --host 127.0.0.1 --port 59999 --mode launchServer',
    getcwd(),
    ['APP_URL' => 'http://127.0.0.1:59999'],
);
$process->setTimeout(0);
$process->start();
$process->waitUntil(fn ($type, $output) => str_contains($output, 'Listening on'));

echo 'symfony getPid(): '.$process->getPid().PHP_EOL;
system("ps -ef --forest | grep -E 'playwright|php' | grep -v grep");

$process->stop(0.1, SIGTERM); // exactly what PlaywrightNpmServer::stop() does

usleep(700000);
system("pgrep -af 'playwright run-server' || echo 'no survivors'");

Observed output on Ubuntu (/bin/sh → dash):

symfony getPid(): 29995
php repro.php
 \_ sh -c ./node_modules/.bin/playwright run-server --host 127.0.0.1 --port 59999 --mode launchServer
 |   \_ node ./node_modules/.bin/playwright run-server --host 127.0.0.1 --port 59999 --mode launchServer
29996 node ./node_modules/.bin/playwright run-server --host 127.0.0.1 --port 59999 --mode launchServer   ← survivor

Suggested fix (verified)

Prepend exec to the shell command in ServerManager::playwright() (non-Windows):

'exec .'.DIRECTORY_SEPARATOR.'node_modules'.DIRECTORY_SEPARATOR.'.bin'.DIRECTORY_SEPARATOR.'playwright run-server --host %s --port %d --mode launchServer',

With exec, sh replaces itself with node, getPid() points at the server itself, and
stop() kills it reliably. Verified with the repro above — the tree collapses to
php → node and no survivors remain after stop(). Also verified end-to-end with a real
browser test run: the terminal returns immediately even with piped output, and
pgrep -af 'playwright run-server' is clean afterwards.

Alternatively, use the array form of Process (no shell at all). Two smaller related
observations spotted while debugging this:

  • PlaywrightNpmServer::stop() passes SIGTERM as the fallback signal of
    Process::stop(), so it never escalates to SIGKILL; passing null (default SIGKILL)
    would be more robust.
  • Plugin::terminate() swallows the Revolt Error («Must call resume() or throw() before
    calling suspend() again») with an early return before stopping the playwright server —
    if that path is ever taken, the server would leak even with the signal fix.

Activity

  1. leonardoaugustus commented on Aug 4, 2026

    @leonardoaugustus

    Confirming this is still present in v5 — the code path is unchanged from the 4.x analysis above.

    Environment

    • pestphp/pest v5.0.3, pestphp/pest-plugin-browser v5.0.0
    • PHP 8.5.8, Laravel Sail runtime on Ubuntu 24.04, /bin/sh → dash
    • playwright 1.61.1, Chromium headless

    ServerManager::playwright() still builds the string command without exec, and PlaywrightNpmServer::start() still runs it through Process::fromShellCommandline(), so the diagnosis carries over verbatim. After a clean, fully passing run:

    $ ./vendor/bin/pest tests/Browser/HelpMenuTest.php
      Tests:    1 passed (7 assertions)
    
    $ ps -eo pid,ppid,etimes,rss,args | grep '[p]laywright run-server'
    72948     1       3 176740 node ./node_modules/.bin/playwright run-server --host 127.0.0.1 --port 37527 --mode launchServer
    

    PPID 1, ~176 MB RSS, one per run. No interruption or failure needed — a green run leaks too.

    A second symptom worth adding: it surfaces as flaky browser tests, not just a hanging terminal.

    We chased this for a while as a flaky test, because the accumulation degrades the machine before it breaks it. Eight consecutive runs of the same 2-test file, measuring after each:

    run orphans container free RAM
    1 1 5753 MB
    2 2 5556 MB
    4 4 5247 MB
    6 6 4969 MB
    8 8 4669 MB

    Linear, ~150 MB per run. Over two days of ordinary development we had 65 orphaned servers and 121 MB free of 7.9 GB. At that point browser interactions start stalling past the plugin's 5 s timeout, and each suite run fails a different test — HelpMenuTest here, ResponsiveLayoutTest there — which reads exactly like flakiness in the tests themselves. Raising the timeout 5 s → 15 s did not help (the failures simply became 15 s timeouts, and runs went from ~20 s to 33–76 s); killing the orphans did, and five consecutive suite runs went green with memory flat.

    So for anyone landing here from a flaky-browser-test search: this is likely your cause, and pkill -f '[p]laywright run-server' before each run is a workable stopgap until the exec fix lands.

    Happy to test a patch on this environment if that helps.

  2. likemusic commented on Sep 22, 2026

    @likemusic
    Author

    Status update, and a pointer to the PR that fixes this.

    Thanks @leonardoaugustus for the v5 confirmation — that matches what is in the tree today. On 5.x, ServerManager still builds the command as ./node_modules/.bin/playwright run-server --host %s --port %d --mode launchServer with no exec prefix (the line is byte-identical to 4.x), and PlaywrightNpmServer::start() still runs it through Process::fromShellCommandline(). One thing did change — #172 ("Stop playwright with default signals") dropped the explicit SIGTERM from stop() — but that does not help here: the signal still reaches the sh wrapper, not node.

    pestphp/pest-plugin-browser#254 fixes exactly this, against 5.x: it adds the exec prefix so node replaces the shell and stop() signals the real process, with a unit test on the composed command line. It supersedes #169 and #211 (both aimed at the dormant 4.x). That PR is the right place to review this.

    Leaving this open until it is merged, since the orphaned node process is still reproducible on a released version.


    Update, 29 September 2026 — confirmed on a third environment, and #254 verified as the fix.

    WSL2 (Ubuntu, /bin/sh → dash), Linux 6.18 kernel on a Windows host, PHP 8.4.24, Playwright 1.62.1, plugin 5.x at c98e8a5. Same test file in both arms, identical outcome per run (1 failed, 26 skipped, 2 passed; the failure is an unrelated external-URL test), counting the servers left behind:

    run stock 5.x with pestphp/pest-plugin-browser#254
    1 1 orphan 0
    2 2 orphans 0
    3 3 orphans 0

    ~182 MB RSS each. On WSL2 the orphans are reparented to /init rather than to PID 1 by name, which is the same thing under a different label — so this environment reads differently in ps while being the same defect.

    The accumulation is as silent here as @leonardoaugustus described: before I went looking, five ordinary runs earlier in the day had left five servers and ~730 MB behind, with every run green and nothing in the output to suggest it.

    One correction to my previous comment: #254 has grown since. As of 28 September it is no longer only the exec prefix — it also keeps body-less responses off keep-alive connections. I verified only the process half, which is what this issue is about.

  3. AlexR1712 commented on Oct 2, 2026

    @AlexR1712

    thank you guys, i hope this will be available soon, it causes a lot of issues when ai agents works for you 😄

  4. sineld commented on Oct 6, 2026

    @sineld

    Another confirmation, on an architecture not covered here yet, plus a data point about parallel runs that I did not see mentioned above.

    Environment

    • pestphp/pest v5.0.2, pestphp/pest-plugin-browser v5.0.0
    • PHP 8.4.25, Ubuntu 24.04 aarch64 (Oracle ARM server, Docker), /bin/sh → dash
    • Playwright 1.62.0, Chromium headless

    Parallel runs leak once per run, not once per worker.

    Full suite, --parallel --processes=8, 620 tests of which 4 are browser tests, all green:

    $ pgrep -cf '[p]laywright run-server'          # before
    0
    $ vendor/bin/pest --parallel --processes=8
      Tests:  618 passed, 2 skipped (620 total)
    $ ps -eo pid,ppid,rss,args | grep '[p]laywright run-server'
    3808309   1   128860   node ./node_modules/.bin/playwright run-server --host 127.0.0.1 --port 55169 --mode launchServer
    

    That lines up with the code: ServerManager::playwright() only builds a PlaywrightNpmServer when Parallel::isWorker() is false, so the one server started in the main process during collection (UsesBrowserTestCaseMethodFilter::accept()) is the only one, and the workers attach to it via AlreadyStartedPlaywrightServer::fromPersisted(). So --parallel neither multiplies nor avoids the leak: Plugin::terminate() does reach playwright()->stop() in the main process, and it still does not kill anything, because the signal lands on the dash wrapper exactly as diagnosed above.

    The orphan also outlives the parent being killed outright. Running the browser file under timeout 600 php vendor/bin/pest tests/Browser/ConsoleFlowTest.php, the PHP process is gone when the limit fires and the server is left at PPID 1 with the same RSS. Same root cause seen from the other side; overlaps #1825.

    Accumulation here: 14 servers / ~1.4 GB RSS over 13 days of ordinary development, every run green and nothing in the output hinting at it. We only went looking because of the memory.

    Confirming the released tag is still affected: PlaywrightNpmServer::start() in v5.0.0 has no exec in the composed command line (grep -n exec on that file returns nothing), so pestphp/pest-plugin-browser#254 is still the fix for the version people are installing today.

    Stopgap in the meantime, for anyone else landing here: pkill -f '[p]laywright run-server' as the first step of the composer test scripts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions