Skip to content

[ci] Track the /flow route bundle size against main - #3739

Merged
VaguelySerious merged 4 commits into
mainfrom
peter/bundle-size-flow-route
Aug 25, 2026
Merged

[ci] Track the /flow route bundle size against main#3739
VaguelySerious merged 4 commits into
mainfrom
peter/bundle-size-flow-route

Conversation

@VaguelySerious

@VaguelySerious VaguelySerious commented Aug 22, 2026

Copy link
Copy Markdown
Member

TLDR adds this comment to each PR that compares output sizes to main

image

Builds the nextjs-turbopack and hono workbench apps on every PR, measures the /.well-known/workflow/v1/flow route, and posts a sticky comment with the delta against main. Pushes to main produce the baseline artifacts PR runs download. Non-required: it lives in its own workflow file, so it cannot gate the E2E aggregate.

Two numbers per app

Neither app emits an isolable function bundle for that route. Next.js emits a ~1 KB turbopack chunk loader pointing at chunks shared with other routes; nitro inlines the handler into a single server entry.

  • Gated: what the workflow builders emit for the flow route, before the framework bundles it. Growth beyond max(2%, 50 KiB) raw fails the job. Override with the allow-bundle-size-growth label.
  • Informational: the framework's own build output. Unrelated changes move it, so it never gates.

Each is only ever compared against its own baseline, never against the other.

Verification

Reproducibility was checked before wiring the gate, since a noisy metric makes one useless: two clean builds of each app produce byte-identical reports.

Flow bundle Step registrations Framework output
nextjs-turbopack 1.32 MiB 1.7 KiB 3.21 MiB
hono 1.28 MiB 261.0 KiB 7.70 MiB

Both loud-failure paths were exercised and exit non-zero: a Tier-1 bundle backdated one day against the build stamp, and a missing _workflows.ts from a skipped prebuild. A plausible wrong number (the 1 KB chunk loader, a stale artifact) is worse than a red job.

Pinned build env

WORKFLOW_SOURCEMAP, WORKFLOW_PUBLIC_MANIFEST and WORKFLOW_TARGET_WORLD are pinned and recorded in each report's fingerprint; the renderer refuses to diff reports whose fingerprints disagree.

WORKFLOW_SOURCEMAP=false is the one that moves the numbers: sourcemap mode defaults to inline outside a production build, which alone takes the Next flow bundle from 1.38 MB to 5.85 MB.

WORKFLOW_TARGET_WORLD was measured to have no effect. Building nextjs-turbopack with local and with vercel gives byte-identical reports on all three metrics, because every world the app depends on is bundled into the framework output either way and the choice is made at runtime. It is pinned anyway, as insurance: packages/next branches on it, so if it ever starts mattering the fingerprint turns that into a refused diff rather than a phantom code change.

Known limitation

The gate does not cover the world adapters. A change confined to @workflow/world-vercel will not move the gated numbers. Documented in the PR comment, the measurement script, and AGENTS.md. Gating world code would need a different metric.

Notes for review

  • The first PR run after this merges will show "No baseline on main yet" and gate nothing. Baselines exist only once this lands on main or the workflow is dispatched there.
  • Per-app bundle paths are read from scripts/create-test-matrix.mjs rather than restated, so there is one source of truth.
  • The app build uses the package script, not turbo run build: hono's turbo.json does not declare node_modules/.nitro/ as an output, so a cache hit would leave the gated file absent.

🤖 Generated with Claude Code

Builds the nextjs-turbopack and hono workbench apps on every PR, measures
the /.well-known/workflow/v1/flow route, and posts a sticky comment with the
delta against main. Pushes to main produce the baseline artifacts.

Two numbers per app, because neither app emits an isolable function bundle
for that route: Next.js emits a ~1 KB turbopack chunk loader pointing at
chunks shared with other routes, and nitro inlines the handler into a single
server entry. The gated number is what the workflow builders emit before the
framework bundles it; the framework's own output is reported but never gates,
since unrelated changes move it.

Verified before wiring the gate: two clean builds of each app produce
byte-identical reports, so a delta means a real change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-bot Bot commented Aug 22, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: b9d0c1b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Aug 25, 2026 2:56pm
example-nextjs-workflow-webpack Ready Ready Preview, v0 Aug 25, 2026 2:56pm
example-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-astro-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-express-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-fastify-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-hono-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-nestjs-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-nitro-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-nuxt-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-python-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-sveltekit-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-tanstack-start-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workbench-vite-workflow Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workflow-docs Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workflow-swc-playground Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workflow-tarballs Ready Ready Preview, v0 Aug 25, 2026 2:56pm
workflow-web Ready Ready Preview, v0 Aug 25, 2026 2:56pm

@github-actions

github-actions Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M0WQ2QFY0GQ8Y7FJ22Z5ZVVV | 🔍 observability
  • sleepingWorkflow | wrun_41M0WQ3CQY0GYZ9D47SVJ55P5R | 🔍 observability
  • parallelSleepWorkflow | wrun_41M0WQ3DHB0GPSEH8VBZBPHNDP | 🔍 observability
  • nullByteWorkflow | wrun_41M0WQ3MSW0GXPHDV3TCWBNXCS | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M0WQ8D1R0GXD2FPZRDVGQTTK | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M0WQ8GW00GHDGC3DBR3WNBTZ | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M0WQ8PM40GT0EMCGFF6XXF5A | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M0WQ98EH0GJWR20V1DHB3W96 | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M0WQ5ZG1Y0QKQCRPWQS6Q7S8
  • promiseAllWorkflow | wrun_41M0WQ2QFY0GQ8Y7FJ22Z5ZVVV
  • sleepingWorkflow | wrun_41M0WQ3CQY0GYZ9D47SVJ55P5R
  • parallelSleepWorkflow | wrun_41M0WQ3DHB0GPSEH8VBZBPHNDP
  • nullByteWorkflow | wrun_41M0WQ3MSW0GXPHDV3TCWBNXCS
  • cancelRun - cancelling a running workflow | wrun_41M0WQ8D1R0GXD2FPZRDVGQTTK
  • cancelRun via CLI - cancelling a running workflow | wrun_41M0WQ8GW00GHDGC3DBR3WNBTZ
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M0WQ8PM40GT0EMCGFF6XXF5A
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M0WQ98EH0GJWR20V1DHB3W96

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • addTenWorkflow (vite)
  • cancelRun via CLI - cancelling a running workflow (express)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • fibonacciWorkflow - recursive workflow composition via start() (nextjs-webpack)
  • health check (CLI) - workflow health command reports healthy endpoints (nextjs-webpack)
  • sleepWinsRaceWorkflow (tanstack-start)
  • webhookWorkflow (fastify)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

34 infra events
  • cold-start-warmup · suite warmup (python) · at 15:01:41Z · abandoned wrun_41M0WPYAE00GM6K1CDJ8DPERHV · (+7 more)
  • run-pickup-stall · promiseAllWorkflow (python) · at 15:01:57Z · abandoned wrun_41M0WQ1Z3D0GSBNMRD14YJH4YH
  • cold-start-warmup · suite warmup (tanstack-start) · at 15:01:57Z · abandoned wrun_01M0WQ1M5P1MRH6VYV0E38KMQK
  • run-pickup-stall · sleepingWorkflow (python) · at 15:01:57Z · abandoned wrun_41M0WQ1Z3G0GRQ9PE5NNA8G2DW
  • run-pickup-stall · nullByteWorkflow (python) · at 15:01:57Z · abandoned wrun_41M0WQ1Z3J0GTXBJ93TZKZGE7Y
  • run-pickup-stall · parallelSleepWorkflow (python) · at 15:01:57Z · abandoned wrun_41M0WQ1Z3G0GRQ9PE5NNA8G2DX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 15:01:57Z · abandoned wrun_41M0WQ1ZE60GJW11XKZMV7XA52
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 15:02:29Z · abandoned wrun_41M0WQ2YX30GMX4WMFK2RBFVCB
  • cold-start-warmup · suite warmup (python) · at 15:02:37Z · abandoned wrun_01M0WQ0129FFFEBEER793NP4MT · (+7 more)
  • run-pickup-stall · promiseAllWorkflow (python) · at 15:02:52Z · abandoned wrun_01M0WQ3P6TZWQ5CYZ54HYTX8WK
  • run-pickup-stall · sleepingWorkflow (python) · at 15:02:52Z · abandoned wrun_01M0WQ3P6ZM2XPN5BN7FHCD7A7
  • run-pickup-stall · parallelSleepWorkflow (python) · at 15:02:52Z · abandoned wrun_01M0WQ3P71C20SWGPRXGX6BWQF
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 15:02:52Z · abandoned wrun_01M0WQ3P6S56GZPJK2YZB4J1PD
  • run-pickup-stall · nullByteWorkflow (python) · at 15:02:52Z · abandoned wrun_01M0WQ3P73WDSRBS47HQWE6T2Z
  • run-pickup-stall · parallelSleepWorkflow (python) · at 15:02:58Z · abandoned wrun_41M0WQ3TEQ0GR0C3ESC9A4H591
  • run-pickup-stall · promiseAllWorkflow (python) · at 15:02:58Z · abandoned wrun_41M0WQ3TH90GM476R7C0VG0GA4
  • run-pickup-stall · sleepingWorkflow (python) · at 15:02:58Z · abandoned wrun_41M0WQ3TJC0GSMEN8BZEY4VXCV
  • run-pickup-stall · nullByteWorkflow (python) · at 15:02:59Z · abandoned wrun_41M0WQ3THE0GRV55VD6YQP50Q7
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 15:03:48Z · abandoned wrun_41M0WQ3Y630GY12BXAV8KC4S1X
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 15:03:52Z · abandoned wrun_01M0WQ5GTSM98T01GVJMRN9DKK
  • run-pickup-stall · parallelSleepWorkflow (python) · at 15:03:52Z · abandoned wrun_01M0WQ5GTYG1NP262A69DJNPX6
  • run-pickup-stall · promiseAllWorkflow (python) · at 15:03:52Z · abandoned wrun_01M0WQ5GTVV4B884BHZQKXMEZW
  • run-pickup-stall · sleepingWorkflow (python) · at 15:03:52Z · abandoned wrun_01M0WQ5GTXCS5X4N6GYS3F2K6D
  • run-pickup-stall · nullByteWorkflow (python) · at 15:03:52Z · abandoned wrun_01M0WQ5GV0P06TKS4MDSQGC65F
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 15:04:52Z · abandoned wrun_01M0WQ7BECN6WHD7BPCXZK3VPV
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 15:04:52Z · abandoned wrun_01M0WQ7BEPJQCFE3WZY252AJC1
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 15:04:52Z · abandoned wrun_01M0WQ7BES58J7ME9WGFJZV8PM
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 15:05:22Z · abandoned wrun_01M0WQ88T3V2QA7XNJJ4QK397W
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 15:05:22Z · abandoned wrun_01M0WQ88T55S08JCJGF7T642ND
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 15:05:29Z · abandoned wrun_41M0WQ5XEJ0GXKC1E0PA52B0NW
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 15:05:30Z · abandoned wrun_41M0WQ5WWT0GQTMT46W0BBZ3A0
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 15:05:52Z · abandoned wrun_01M0WQ961YBTQSJ78TRA3G9VB2
  • run-pickup-stall · StepNotRegisteredError fails the step but workflow can catch it (nextjs-webpack) · at 15:06:30Z · abandoned wrun_01M0WQABCRB90A05JYXWEK292Q
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 15:06:48Z · abandoned wrun_01M0WQAWGKMNZ0W8C74Z5CFNXY

E2E Test Summary

Summary
Passed Failed Skipped Total
❌ ▲ Vercel Production 3570 8 742 4320
✅ 💻 Local Development 3922 0 558 4480
✅ 📦 Local Production 3922 0 558 4480
✅ 🐘 Local Postgres 3922 0 558 4480
✅ 🪟 Windows 320 0 0 320
❌ 🌐 Cross-language Conformance 0 9 132 141
✅ vercel-http-transport 817 0 143 960
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 553 0 87 640
Total 17053 17 2778 19848
Details by Category

❌ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 132 0 28
✅ astro-quickjs 132 0 28
✅ example-node 132 0 28
✅ example-quickjs 132 0 28
✅ express-node 132 0 28
✅ express-quickjs 132 0 28
✅ fastify-node 132 0 28
✅ fastify-quickjs 132 0 28
✅ hono-node 132 0 28
✅ hono-quickjs 132 0 28
✅ nest-node 132 0 28
✅ nest-quickjs 132 0 28
✅ nextjs-turbopack-node 157 0 3
✅ nextjs-turbopack-quickjs 157 0 3
✅ nextjs-webpack-node 157 0 3
✅ nextjs-webpack-quickjs 157 0 3
✅ nitro-node 132 0 28
✅ nitro-quickjs 132 0 28
✅ nuxt-node 132 0 28
✅ nuxt-quickjs 132 0 28
❌ python-node 0 8 152
✅ sveltekit-node 151 0 9
✅ sveltekit-quickjs 151 0 9
✅ tanstack-start-node 132 0 28
✅ tanstack-start-quickjs 132 0 28
✅ vite-node 132 0 28
✅ vite-quickjs 132 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 0
✅ nextjs-turbopack-quickjs 160 0 0

❌ 🌐 Cross-language Conformance

App Passed Failed Skipped
❌ python 0 9 132

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ hono 132 0 28
✅ nextjs-turbopack 157 0 3
✅ nitro 132 0 28
✅ vite 132 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ nextjs-turbopack 157 0 3
✅ vite 132 0 28

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit b9d0c1b · Tue, 25 Aug 2026 15:18:15 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1336 (+62%) 🔻 1461 🔴 (+30%) 🔻 1510 🔴 (+29%) 🔻 1704 🔴 (+15%) 30
TTFS stream 999 (+350%) 🔻 1394 🔴 (+24%) 🔻 1416 🔴 (+25%) 🔻 1450 🔴 (-1.7%) 30
TTFS hook + stream 550 (-56%) 💚 1788 🔴 (+30%) 🔻 1873 🔴 (+22%) 🔻 2008 🔴 (+20%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 660 (+6.5%) 2149 (+28%) 🔻 2300 (+32%) 🔻 2440 (-54%) 💚 10
Fan-out TTLS Promise.all(100 steps) 2112 (+10%) 4222 (+14%) 4584 (+12%) 10550 (-4.6%) 10
STSO 1020 steps (inline) 136 (+62%) 🔻 185 (+43%) 🔻 219 (+44%) 🔻 315 (+25%) 🔻 1019
WO 1020 steps 183322 (+40%) 🔻 183322 (+40%) 🔻 183322 (+40%) 🔻 183322 (+40%) 🔻 1
CRTT first chunk (pooled) 104 (+14%) 163 (+12%) 244 (+11%) 324 (-46%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 142 (+23%) 154 (+2%) 235 (-35%) 502 (-27%) 153 (-25%) 10
size sweep (100/s, 160B-12KB) 131 (+12%) 168 (-22%) 264 (-30%) 400 (-46%) 119 (-49%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 198 (+100%) 129 (+9%) 171 (+13%) 322 (-17%) 162 (-50%) 3
replay eve-gpt-5.6-sol-2000t (1x) 165 (+37%) 175 (+28%) 231 (+5%) 1174 (+40%) 571 (-1%) 2
replay eve-gpt-5.6-sol-2000t (2x) 151 (-7%) 222 (-6%) 368 (-14%) 1031 (+44%) 283 (+20%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 130759ms → this run 183012ms (Δ +52253ms, +40%)

 50-100 ms  ┃                         main   1  this   0    -1
100-150 ms  ┃███████████████████████  main 900  this  46  -854
150-200 ms  ██░░░░░░░░░░░░░░░░░░░┃    main  91  this 814  +723
200-250 ms  █░┃                       main  16  this 104   +88
250-300 ms  ┃                         main   8  this  41   +33
300-350 ms  ┃                         main   2  this  10    +8
350-400 ms  ┃                         main   1  this   3    +2
650-700 ms  ┃                         main   0  this   1    +1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90          p99     n
control  ······▃█▁▁···   133.7 (+2%)  123 (+23%)  235 (-35%)   502 (-27%)  3000
sweep    ······▃█▁····   136.6 (-6%)   124 (+8%)  264 (-30%)   400 (-46%)  3000
gw 1x    ·····▁▇█▁····   110.8 (+6%)   103 (+6%)  171 (+13%)   322 (-17%)  5295
eve 1x   ·····▁▂█▂▁▁··  167.4 (+35%)  136 (+42%)   231 (+5%)  1174 (+40%)  5186
eve 2x   ·····▁▂█▃▁▁··   179.1 (+5%)   143 (+1%)  368 (-14%)  1031 (+44%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▃▃▂▇▃▄▂▁▂█  116–162ms
sweep    ▆▁▃▆▃▄▆▄▅█  124–149ms
gw 1x    ▄█▃▄▇▃▁▂▄▄  101–123ms
eve 1x   █▁▁▂▂▂▁▂▁▁  131–354ms
eve 2x   █▂▂▂▁▃▃▄▄▂  128–297ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▆██▅▃▃▁  135–138ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▁▃▆█▄▇▃▃▄▄  29–47ms
sweep    ▁▅▆█▅▆▆▇▅▅  40–59ms
gw 1x    ▁█▅▂▅▃▄▅▃▃  27–35ms
eve 1x   █▁▁▂▄▅▄▃▂▂  22–36ms
eve 2x   █▅▄▂▄▂▃▂▁▃  18–32ms
📜 Previous results (2)

de8da08

Mon, 24 Aug 2026 19:22:14 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 288 (-71%) 💚 1285 🔴 (+16%) 🔻 1304 🔴 (+13%) 1403 🔴 (+12%) 30
TTFS stream 1245 (+76%) 🔻 1306 🔴 (+19%) 🔻 1335 🔴 (+20%) 🔻 1797 🔴 (+40%) 🔻 30
TTFS hook + stream 1522 (+11%) 1622 🔴 (+10%) 1666 🔴 (+11%) 1824 🔴 (+14%) 30
Fan-out TTFS Promise.all(100 steps) 667 (+2.6%) 934 (-45%) 💚 949 (-47%) 💚 1693 (-13%) 10
Fan-out TTLS Promise.all(100 steps) 1854 (-4.0%) 2508 (-21%) 💚 2620 (-20%) 💚 4776 (-44%) 💚 10
STSO 1020 steps (inline) 122 (+16%) 🔻 163 (+10%) 183 (+8.9%) 280 (+28%) 🔻 1019
WO 1020 steps 163721 (+11%) 163721 (+11%) 163721 (+11%) 163721 (+11%) 1
CRTT first chunk (pooled) 88 (±0%) 156 (+26%) 🔻 189 (-2.1%) 271 (-56%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 113 (±0%) 128 (-2%) 195 (-8%) 373 (-51%) 122 (-15%) 10
size sweep (100/s, 160B-12KB) 108 (-4%) 134 (-6%) 174 (-7%) 431 (-62%) 118 (-14%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 120 (-2%) 120 (-82%) 154 (-95%) 266 (-94%) 174 (-20%) 3
replay eve-gpt-5.6-sol-2000t (1x) 128 (+8%) 116 (-13%) 145 (-14%) 321 (-7%) 272 (+4%) 2
replay eve-gpt-5.6-sol-2000t (2x) 156 (+27%) 157 (-14%) 222 (-16%) 727 (+32%) 388 (+70%) 3

c9af771

Sat, 22 Aug 2026 00:57:36 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 268 (-74%) 💚 1331 🔴 (+16%) 🔻 1371 🔴 (+16%) 🔻 1545 🔴 (+18%) 🔻 30
TTFS stream 240 (-77%) 💚 1373 🔴 (+21%) 🔻 1409 🔴 (+21%) 🔻 1515 🔴 (+26%) 🔻 30
TTFS hook + stream 1534 (+26%) 🔻 1723 🔴 (+21%) 🔻 1753 🔴 (+17%) 🔻 1975 🔴 (+2.4%) 30
Fan-out TTFS Promise.all(100 steps) 539 (-28%) 💚 2016 (+4.3%) 2044 (+5.1%) 2069 (+1.9%) 10
Fan-out TTLS Promise.all(100 steps) 1709 (-66%) 💚 5864 (-21%) 💚 7006 (-11%) 7951 (-7.5%) 10
STSO 1020 steps (inline) 136 (+31%) 🔻 194 (+37%) 🔻 222 (+38%) 🔻 337 (+73%) 🔻 1019
WO 1020 steps 189480 (+33%) 🔻 189480 (+33%) 🔻 189480 (+33%) 🔻 189480 (+33%) 🔻 1
CRTT first chunk (pooled) 95 (-22%) 💚 157 (-11%) 184 (-16%) 💚 281 (-49%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 126 (-8%) 161 (-65%) 239 (-63%) 372 (-93%) 133 (-14%) 10
size sweep (100/s, 160B-12KB) 139 (-15%) 179 (-10%) 278 (-9%) 4827 (+963%) 170 (-9%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 152 (-15%) 138 (-47%) 187 (-68%) 612 (-58%) 296 (-64%) 3
replay eve-gpt-5.6-sol-2000t (1x) 145 (-4%) 125 (-34%) 168 (-51%) 293 (-78%) 239 (-70%) 2
replay eve-gpt-5.6-sol-2000t (2x) 147 (-16%) 192 (-29%) 344 (-9%) 1005 (+41%) 412 (+11%) 3
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 211.3 KiB 40.1 KiB 1.76 MiB
nextjs-turbopack 216.7 KiB 439 B 771.9 KiB
About these numbers

Sizes are gzip.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
No baseline on main yet for some rows.

b9d0c1b · run

The same hono build produces 55 files under .output/server on a Linux runner
and 54 on macOS, so a report measured off-runner is not comparable to one
measured on it. Node's major version was already recorded because zlib ships
with Node and moves the gzip numbers; platform and arch cover the raw ones.

Cheap to add now: no baseline exists on main yet, so nothing is invalidated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The table is the comment; the notes explaining what the numbers are, what
gates the job, and which commit produced them sit behind a disclosure so
the default view is four rows and nothing else.
@VaguelySerious
VaguelySerious force-pushed the peter/bundle-size-flow-route branch from 2bafe9c to b9d0c1b Compare August 25, 2026 14:52

@alangenfeld alangenfeld left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice seems like good defense to have

@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for c8bcde5 (AI decision).

This commit adds a brand-new CI capability (a bundle-size measurement workflow, measurement/rendering scripts, tests, and AGENTS.md documentation for it) rather than fixing anything broken on stable. It is not needed to keep the maintenance line buildable, testable, or releasable, and its triggers are scoped to main anyway, so it would gate nothing useful on stable.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

c8bcde53d01781454bdb2fbef05577d843f91aae

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants