Purpose: give an AI assistant a lightweight, evidence-based way to prepare work for Arena Agent Mode and review the result. Arena is a supplementary research/coding executor, not an automated API backend.
Important: Arena Terms §5 restrict automated/programmatic access to the Arena service. A delegating assistant may prepare a GitHub task brief and later review its PR, but the user opens Arena, connects the selected repo and submits the task manually. Do not script Arena's UI, scrape credits or invent an Agent Mode API. Read the data policy before linking a repository.
| Your role | Read first | Responsibility |
|---|---|---|
| Claude, ChatGPT, Codex, Gemini, etc. preparing and reviewing a task | DELEGATE.md | Choose bounded work, write a brief, give the user one launch instruction, review actual PR/diff/tests. |
| Arena Agent executing a task in its connected repo | EXECUTOR.md | Check actual branch/permissions, run the work, deliver a short result file and PR. |
| Exploring actual Arena capabilities | AGENTS.md → capabilities | Distinguish observed operations from documentation and untested claims. |
- Delegator writes
.arena/tasks/0007.md(or an issue) in the target project repo, following templates/task.md. - User opens Arena Agent Mode, connects that same project repo and sends the one-line launch instruction.
- Arena reads the brief, performs bounded work on the actual assigned branch, runs checks and delivers a PR and
.arena/results/0007.md. - Delegator reads the result, actual diff and test evidence, and tells the user whether acceptance criteria were met. User decides whether to merge.
No step requires programmatic access to Arena itself. GitHub is the coordination and artifact channel.
- GitHub commit/push/PR, native public research, npm/pip, Python/Node and local servers were demonstrated.
- Local Chromium checks were demonstrated both on a synthetic fixture and on the real generated Design Spells static build (PR #43). Its
npm run buildexecutespython3 scripts/build.py, and the existingnpm testsuite was not run in PR #43; this was not proof of an Astro runtime or an Astro CLI build. - A 14-source research exercise with a 36-claim ledger is in PR #6; those verification counts are the agent's recorded checks, not an independent accuracy score.
- Design Spells PR #43 recorded
--no-sandboxand--disable-web-securityin the effective launch args but did not commit the launch script, so the latter flag's origin is unverified. Playwright normally adds--no-sandboxunlesschromiumSandbox: true. See browser caveat.
Not established: unrestricted public browser access, exact Agent Mode quotas, a fixed model identity, customer-data confidentiality, or full Lighthouse/WCAG/CWV compliance. Read limitations and modes.
Reusable session inventory — an optional, read-only-by-default snapshot of the current sandbox and exposed agent tools. Run only when a task needs fresh capability evidence; do not treat historical observations as current-session guarantees. Package installation probes require explicit opt-in.
Capability matrix · GitHub · Research · Browser · Data handling · Benchmark · Lightweight job ledger · Historical evidence
Last curated: 2026-09-26. This is an independent field guide, not official product documentation.