Give lab_setup fragments somewhere to run - #108
Merged
Merged
Conversation
`lab_setup` is the one part of a bank entry the template had no mechanism for. Two of the four entries ship one, both load-bearing, and in a deployment built from the template the fragment could not run or even be read: ./lab is a build context, the Dockerfile COPYed nothing, and `command: sleep infinity` would not have sourced it if it had. So `github` and `gcp` installed cleanly and did not work. `github`'s fragment wires the credential helper and sets credential.useHttpPath false, without which the gateway's exact-match route cannot satisfy the per-path lookups git makes; `gcp`'s writes the inert ADC file that lets a Google client library's chain reach a token exchange the proxy can answer, without which it raises before opening a socket. The mechanism is an entrypoint baked into the lab image, running every *.sh in /etc/agent-setup.d in filename order before exec'ing the command. In the image rather than the deployment because a fragment that nothing runs is not an installed provider, and a deployment should not have to supply that itself. Both lab Dockerfiles carry it, kept identical the way the existing copy-diff lint already requires. Executed, not sourced, though both fragments' headers claimed otherwise. Every effect they have is a file they write and neither exports anything, so a child process is enough — and sourcing would put each fragment's `set -euo pipefail` into the entrypoint's shell and let a stray `exit` kill it. It would buy nothing even for a fragment that wanted it: `docker exec` builds its environment from the container's config, not from PID 1, so an export here never reaches an agent that shells in. The headers now say executed, and bank/gcp's fragment goes 644 -> 755 so the two entries stop differing over a bit the entrypoint deliberately does not depend on. The mount is ./lab/setup.d, not ./lab as first sketched: ./lab is the build context and now holds entrypoint.sh, which a *.sh glob would pick up and run inside itself. Read-only, because a fragment can already do anything the lab can do — that is why an entry may ship shell — but persisting a change back into the deployment is a different thing. On start rather than on create, so adding a provider is a file plus a restart. One file per provider, named for it, because the path inside an entry is fixed and two would collide on setup.sh — which makes uninstall `rm` rather than finding and unpicking lines in a shared file. A failing fragment stops the container. A lab that comes up with a provider installed and unconfigured is the failure this exists to prevent, so it is not logged past. 80-lab-setup covers the behaviour on a small base image, the real one being 1.8GB that this tier will not pull; 00-config-lint asserts the part that test cannot see, that both Dockerfiles actually COPY the script and run it. Also covered: the empty-directory case, where `printf '%s\n' dir/*.sh` under nullglob emits one empty line and an empty directory reads as a fragment named "" — the path most deployments take, since most install no lab_setup provider. PLAYBOOK step 7 previously said to "make sure your setup.sh runs it", naming a file the template does not ship; it now says where the fragment goes and that nothing else needs wiring. The schema's lab_setup description said "on create" and said nothing about how it is invoked. Closes #107 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HZotg9g1oPFfkJjGGQLRhJ
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #107. The gap is exactly as reported; two things in the framing are worth correcting, and one detail of the suggested shape doesn't survive contact.
What was actually wrong
lab_setupis the one part of a bank entrytemplate/deployment/had no mechanism for../labis a build context, the Dockerfile COPYed nothing, andcommand: sleep infinitywouldn't have run a fragment if one had been there. Sogithubandgcpinstalled cleanly and did not work.Correction: this is not a boundary failure
The issue's strongest claim is that
ghwithoutgit_protocol https"quietly stops being mediated". Under the default it doesn't — it fails closed. Thelabnetwork shipsinternal: true, so there is no default gateway and SSH has no route out at all. The repo already says so in both places that matter:template/deployment/compose.yaml:249namessshas the reason to setLAB_INTERNAL=false, andPLAYBOOK.md:275-279says the symptom of needing that opt-out is DNS failure inside the lab.The claim holds only under
LAB_INTERNAL=false— where egress mediation is knowingly off,check-invariants.shreports it every run, andghis not special because everything bypasses.That changes the severity, not the fix. What's true is that two of four bank entries install non-functional, and nothing says so.
Correction: the mount can't be
./lab./labis the build context and now holdsentrypoint.sh— a*.shglob there would pick up the entrypoint and run it inside itself. Fragments live in./lab/setup.d/, mounted at/etc/agent-setup.d, which also mirrorscred-gateway/gateway.d/.Executed, not sourced
Both fragments' headers claimed "Sourced by the deployment's
setup.sh". They shouldn't be, and don't need to be — every effect either one has is a file it writes (git config --global, the inert ADC document), and neither exports anything.Sourcing would put each fragment's
set -euo pipefailinto the entrypoint's shell and let a strayexitkill it. And it would buy nothing even for a fragment that wanted it:docker execbuilds its environment from the container's config, not from PID 1, so an export from an entrypoint never reaches an agent that shells in. A fragment needing an agent-visible variable writes/etc/profile.d/, or the entry declareslab_env.Consequences: headers corrected, and
bank/gcp/lab/setup.shgoes 644 → 755 so the two entries stop differing over a bit the entrypoint deliberately doesn't depend on (it invokesbashexplicitly, and a mount can flatten the bit anyway).Shape
lab/setup.d/<name>.sh— the path inside an entry is fixed (lab/setup.sh), so two would collide. This is what makes uninstallrmrather than finding and unpicking lines in a shared file, which is the fragile version of "hooks".Tests
New
tests/integration/80-lab-setup.test.sh, 18 assertions. It runs the realentrypoint.shin the real container shape on a small base — the actual lab image istypescript-node:22at 1.8GB, which this tier won't pull, and which is whytests/stacks/20-boundaryalready refuses to start the lab at all.00-config-lintasserts the part that test can't see: that both Dockerfiles genuinelyCOPYthe script and setENTRYPOINT/CMD, so the mechanism can't be deleted from the image with every behavioural test still green.One case worth calling out, because it's the path most deployments take: an empty
setup.d/.printf '%s\n' dir/*.shundernullglobstill emits one empty line, so an empty directory read as a single fragment named""and refused to start. Caught while testing, fixed, and covered.Local runs of all four CI jobs: integration 17 suites green,
stacks 10andstacks 20green.Targets
release/1.14.0as a feature.🤖 Generated with Claude Code
https://claude.ai/code/session_01HZotg9g1oPFfkJjGGQLRhJ