Skip to content

Give lab_setup fragments somewhere to run - #108

Merged
lpezet merged 1 commit into
release/1.14.0from
feature/lab-setup-hooks
Aug 17, 2026
Merged

lpezet merged 1 commit into
release/1.14.0from
feature/lab-setup-hooks

Conversation

@lpezet

@lpezet lpezet commented Aug 17, 2026

Copy link
Copy Markdown
Owner

Closes #107. The gap is exactly as reported; two things in the framing are worth correcting, and one detail of the suggested shape doesn't survive contact.

What was actually wrong

lab_setup is the one part of a bank entry template/deployment/ had no mechanism for. ./lab is a build context, the Dockerfile COPYed nothing, and command: sleep infinity wouldn't have run a fragment if one had been there. So github and gcp installed cleanly and did not work.

Correction: this is not a boundary failure

The issue's strongest claim is that gh without git_protocol https "quietly stops being mediated". Under the default it doesn't — it fails closed. The lab network ships internal: true, so there is no default gateway and SSH has no route out at all. The repo already says so in both places that matter: template/deployment/compose.yaml:249 names ssh as the reason to set LAB_INTERNAL=false, and PLAYBOOK.md:275-279 says the symptom of needing that opt-out is DNS failure inside the lab.

The claim holds only under LAB_INTERNAL=false — where egress mediation is knowingly off, check-invariants.sh reports it every run, and gh is not special because everything bypasses.

That changes the severity, not the fix. What's true is that two of four bank entries install non-functional, and nothing says so.

Correction: the mount can't be ./lab

./lab is the build context and now holds entrypoint.sh — a *.sh glob there would pick up the entrypoint and run it inside itself. Fragments live in ./lab/setup.d/, mounted at /etc/agent-setup.d, which also mirrors cred-gateway/gateway.d/.

Executed, not sourced

Both fragments' headers claimed "Sourced by the deployment's setup.sh". They shouldn't be, and don't need to be — every effect either one has is a file it writes (git config --global, the inert ADC document), and neither exports anything.

Sourcing would put each fragment's set -euo pipefail into the entrypoint's shell and let a stray exit kill it. And it would buy nothing even for a fragment that wanted it: docker exec builds its environment from the container's config, not from PID 1, so an export from an entrypoint never reaches an agent that shells in. A fragment needing an agent-visible variable writes /etc/profile.d/, or the entry declares lab_env.

Consequences: headers corrected, and bank/gcp/lab/setup.sh goes 644 → 755 so the two entries stop differing over a bit the entrypoint deliberately doesn't depend on (it invokes bash explicitly, and a mount can flatten the bit anyway).

Shape

  • In the image, both lab Dockerfiles, kept identical as the existing copy-diff lint requires. A fragment nothing runs is not an installed provider, and a deployment shouldn't have to supply that mechanism itself.
  • One file per provider, lab/setup.d/<name>.sh — the path inside an entry is fixed (lab/setup.sh), so two would collide. This is what makes uninstall rm rather than finding and unpicking lines in a shared file, which is the fragile version of "hooks".
  • On start, not on create — adding a provider is a file plus a restart, never a rebuild.
  • Read-only mount. A fragment can already do anything the lab can do; persisting a change back into the deployment is a different thing.
  • A failing fragment stops the container, rather than coming up half-configured.

Tests

New tests/integration/80-lab-setup.test.sh, 18 assertions. It runs the real entrypoint.sh in the real container shape on a small base — the actual lab image is typescript-node:22 at 1.8GB, which this tier won't pull, and which is why tests/stacks/20-boundary already refuses to start the lab at all. 00-config-lint asserts the part that test can't see: that both Dockerfiles genuinely COPY the script and set ENTRYPOINT/CMD, so the mechanism can't be deleted from the image with every behavioural test still green.

One case worth calling out, because it's the path most deployments take: an empty setup.d/. printf '%s\n' dir/*.sh under nullglob still emits one empty line, so an empty directory read as a single fragment named "" and refused to start. Caught while testing, fixed, and covered.

Local runs of all four CI jobs: integration 17 suites green, stacks 10 and stacks 20 green.

Targets release/1.14.0 as a feature.

🤖 Generated with Claude Code

https://claude.ai/code/session_01HZotg9g1oPFfkJjGGQLRhJ

`lab_setup` is the one part of a bank entry the template had no mechanism
for. Two of the four entries ship one, both load-bearing, and in a deployment
built from the template the fragment could not run or even be read: ./lab is a
build context, the Dockerfile COPYed nothing, and `command: sleep infinity`
would not have sourced it if it had.

So `github` and `gcp` installed cleanly and did not work. `github`'s fragment
wires the credential helper and sets credential.useHttpPath false, without
which the gateway's exact-match route cannot satisfy the per-path lookups git
makes; `gcp`'s writes the inert ADC file that lets a Google client library's
chain reach a token exchange the proxy can answer, without which it raises
before opening a socket.

The mechanism is an entrypoint baked into the lab image, running every *.sh in
/etc/agent-setup.d in filename order before exec'ing the command. In the image
rather than the deployment because a fragment that nothing runs is not an
installed provider, and a deployment should not have to supply that itself.
Both lab Dockerfiles carry it, kept identical the way the existing copy-diff
lint already requires.

Executed, not sourced, though both fragments' headers claimed otherwise. Every
effect they have is a file they write and neither exports anything, so a child
process is enough — and sourcing would put each fragment's `set -euo pipefail`
into the entrypoint's shell and let a stray `exit` kill it. It would buy
nothing even for a fragment that wanted it: `docker exec` builds its
environment from the container's config, not from PID 1, so an export here
never reaches an agent that shells in. The headers now say executed, and
bank/gcp's fragment goes 644 -> 755 so the two entries stop differing over a
bit the entrypoint deliberately does not depend on.

The mount is ./lab/setup.d, not ./lab as first sketched: ./lab is the build
context and now holds entrypoint.sh, which a *.sh glob would pick up and run
inside itself. Read-only, because a fragment can already do anything the lab
can do — that is why an entry may ship shell — but persisting a change back
into the deployment is a different thing.

On start rather than on create, so adding a provider is a file plus a restart.
One file per provider, named for it, because the path inside an entry is fixed
and two would collide on setup.sh — which makes uninstall `rm` rather than
finding and unpicking lines in a shared file.

A failing fragment stops the container. A lab that comes up with a provider
installed and unconfigured is the failure this exists to prevent, so it is not
logged past.

80-lab-setup covers the behaviour on a small base image, the real one being
1.8GB that this tier will not pull; 00-config-lint asserts the part that test
cannot see, that both Dockerfiles actually COPY the script and run it. Also
covered: the empty-directory case, where `printf '%s\n' dir/*.sh` under
nullglob emits one empty line and an empty directory reads as a fragment named
"" — the path most deployments take, since most install no lab_setup provider.

PLAYBOOK step 7 previously said to "make sure your setup.sh runs it", naming a
file the template does not ship; it now says where the fragment goes and that
nothing else needs wiring. The schema's lab_setup description said "on create"
and said nothing about how it is invoked.

Closes #107

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HZotg9g1oPFfkJjGGQLRhJ
@lpezet
lpezet merged commit 84df42b into release/1.14.0 Aug 17, 2026
4 checks passed
@lpezet lpezet mentioned this pull request Aug 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant