Skip to content

Refactor CI test sharding: Dynamic bin-packing to replace hardcoded suites #40

Description

@hongwei1

Problem

The current test sharding strategy across CI workflows (build_pull_request.yml, build_container.yml) and the local test runner (run_tests_parallel.sh) is based on hardcoded, statically balanced package lists with a fallback catch-all.

While sharding is necessary (to parallelize the heavy 40+ minute single-job integration test load across 9 VMs), the current implementation has several severe drawbacks:

  1. Static drift and Coverage Black Holes: Statically defined packages easily drift. We recently discovered and fixed a coverage black hole where newly added suites fell through the cracks of the hand-crafted inclusions and the catch-all.
  2. Three Sources of Truth: The shard lists are duplicated across three different files (build_pull_request.yml, build_container.yml, and run_tests_parallel.sh).
  3. Fragile Catch-all: The catch-all mechanism has package-level vs class-level semantic traps.
  4. Manual Rebalancing: As tests grow, shards become unbalanced (e.g., historical splits between v6 and v2_x), requiring manual intervention.

Proposed Solution

Instead of maintaining three hardcoded lists, we should introduce a single dynamic script (.github/scripts/compute_shards.py) to handle the shard distribution automatically using timing-based bin-packing.

The script will:

  1. Enumerate all suite classes dynamically to guarantee mathematically complete coverage (eliminating the need for a catch-all).
  2. Read previous timing data (which is already generated per-class by the report jobs) from a cache or artifact.
  3. Bin-pack the suites into N evenly distributed shards using a greedy algorithm.
  4. Output a matrix JSON string.

Both GitHub Action workflows can use jobs.<id>.outputs + fromJSON() to dynamically generate their matrix, and the local run_tests_parallel.sh can call the same script to distribute local shards.

This unifies the shard generation into a single source of truth, guarantees 100% test coverage by construction, and ensures optimal CI time without manual rebalancing.

Activity

  1. hongwei1 commented on Jul 3, 2026

    @hongwei1
    OwnerAuthor

    One more thing worth capturing for later, separate from the dynamic bin-packing fix above (that one should land regardless):

    Root cause behind needing 9 VMs at all: of the ~2921 tests, ~2477 are integration (boot a real server, hit H2). forkMode=once means the whole suite shares one JVM, one embedded server, one DB, one connection pool — so parallelism is only achievable by forking separate OS processes/VMs, never within a single JVM. Bin-packing (this issue) optimizes how those 9 shards are assigned; it doesn't reduce why 9 are needed.

    The structural fix would be per-fork isolation (dedicated port + DB per parallel fork within one JVM/process), which would let far more of the 2477 integration tests run concurrently on fewer, cheaper runners. That's a much bigger change than shard bin-packing — flagging it here as a possible future issue, not something to take on as part of this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions