Shahzad Bhatti Welcome to my ramblings and rants!

September 28, 2026

Governing the AI Factory: How to Ship Fast Without Derailing the Release Train

Filed under: Computing — admin @ 4:51 pm

Here’s how to define where AI fits, reduce cognitive and intent debt, and keep your delivery pipeline stable.


1. The AI Genie

AI coding agents can produce code at an incredible speed but without the quality, you will find that it quickly derails the release train. The existing CI and code-review processes were designed for human-paced changes but this doesn’t keep up when you are pushing code at 10x and you end up with the release problem.

Here is how some of the companies are trying to solve merge queue problems:

SourceFinding
LinearCut PR wait time via faster runners, tsc -> tsgo, sparse checkout, additional test shards
Atlassian70+ large repos on merge queues, PR-level CI answers “does this work alone,” not “does this still work with everything”
NirvanaStateless speculative merge queue.
MergifySpeculative testing, batching, and scope-aware parallel lanes
AutonomaParallel AI subagents multiply conflict surface area unless generation is batched and merges serialized

Here is what I found how AI is affecting the quality:

SourceFinding
DORA 2025PRs merged per developer +98%. Incidents per PR +243%. Bugs per developer +54%
Cortex 2026PRs per author +20% YoY. Incidents per PR +23.5%. Change failure rate ~30%
GitHub Octoverse 202492% of developers use AI coding tools. GitHub Actions CI/CD minutes up 169%
GitClear (211M lines)Code churn doubled. Refactored code fell 24.1% ? 9.5%. Duplicated blocks rose 8x
Uplevel (800 devs, 3 months)+41% bugs after Copilot adoption. No improvement in cycle time or throughput
Opsera (250k+ devs, 60+ orgs)AI PRs sit 4.6x longer in review queues; duplication 10.5% ? 13.5%

Joe Magerramov’s post shows how to build a small Monte Carlo simulator for modeling merge queues. Joe showed how CI/CD becomes a traffic jam, not just a queue. CI/CD batches are cumulative so one defect forces a revert and next batch is affected. I built my own simulator based on Joe’s model (see https://github.com/bhatti/simulators) with additional support for the batched PRs release. It models defect rate with the pipeline duration at 100 commits per day:

The key lesson is that you need to either lower the defect rate per commit or shorten pipeline duration. This is not an easy task, I have encountered a large pipeline duration at many organizations due to large mono-repos, a large codebase with millions of LOC, and large test suite with a long vaidation cycle. The AI agents makes it worse with parallel changes that might conflict resulting in cognitive and intent debt for engineers because no one can track all changes. I explained some of these concepts in my earlier blog and showed how learning feedback loops can be used to build resilient software factories. In order to build end to end agentic SDLC process, you need to define what is AI responsible for and what are the roles for humans.

QuadrantRole
Human Real-timeDecisions that require judgment, context, accountability
Human AsyncReview that needs thought but not immediacy
AI Real-timeAssistance that augments human work in the moment
AI BackgroundAutonomous work that runs without blocking humans

2. The Architecture

I built an orchestration engine Formicary to create data pipelines and CI/CD processes many years ago but I have been using it for driving AI driven workflows. It defines simple primitives to build DAG tasks with exit-code routing, artifact handoff, and fan-out. Here are a few approaches that I am using with the AI driven workflows:

  • Speculative merge queue (test PR N as if N-1 already merged)
  • Scope-aware parallel lanes
  • Incremental builds, test-impact analysis, sparse checkout
  • Deterministic gates a green build can’t waive
  • Independent AI review
  • Contract testing + canary
  • Formal verification of queue invariants (TLA+, Dafny)
  • API fuzz testing
  • Learning flywheel: merge -> extract learnings -> next run reads them -> audit proposes skill changes

Here is how I use Formicary with a CLI toolkit and skillsThe system is split across three repos, each with a clear responsibility boundary:

Here are the core design principles

  • Declarative DAG with a task block that uses on_exit_code for routing
  • Harness + sandbox + skills as separate layers
  • File-based state handoff
  • The learning flywheel closes the loop on review time and merge-time

2.1 Scope router: ai-scope-router

Runs right after create-pr in the existing pipeline and computes a scope key from touched paths. The backing script (scripts/mq/scope_router.py) computes blast radius from line counts and module count, and labels the PR.

# ai-scope-router.yaml (excerpt — full file in docs/examples/)
job_type: ai-scope-router
max_concurrency: 20
timeout: 600s

tasks:
- task_type: classify
  method: KUBERNETES
  script:
    - python -m scripts.mq.scope_router --pr-number {{.PRNumber}}
    - python -m scripts.mq.risk_score --pr-number {{.PRNumber}}
  on_completed: route

- task_type: route
  script:
    - |
      python3 -c "
      import json, os
      scope = json.load(open('/workspace/scope.json'))
      risk = json.load(open('/workspace/risk_score.json'))
      decision = {
          'scope_key': scope['scope'],
          'risk_tier': risk['tier'],
          'lane': scope['scope'],
          'requires_approval': risk.get('requires_human_approval', False)
      }
      json.dump(decision, open('/workspace/route_decision.json', 'w'), indent=2)
      "
  on_completed: done

Here’s what the Slack report looks like for a PR:

# PR Review Report — PR #4091

## Review Findings

? No issues found

## Risk Score

? MEDIUM (score 25.5/100) — Standard review needed; test independently before merge

| Dimension       | Score | Weight | Evidence                                          |
|-----------------|-------|--------|---------------------------------------------------|
| Size            | 7/10  | 1.5×   | 219 lines (+147/?72)                              |
| File Count      | 2/10  | 1.0×   | 5 files changed                                   |
| Blast Radius    | 5/10  | 2.0×   | blast=medium, scope=payments-service              |
| Sensitive Paths | 0/10  | 2.5×   | no sensitive files detected                       |
| Test Coverage   | 0/10  | 1.5×   | good coverage (test:source ?1:1)                  |
| Historical      | 3/10  | 1.0×   | ?? no defect history available — neutral default   |



## Scope

| Field          | Value            | Description                                                    |
|----------------|------------------|----------------------------------------------------------------|
| Scope          | payments-service | All changes owned by payments-service — can merge in dedicated lane |
| Blast radius   | medium           | Moderate change (51–300 lines or 2 modules) — test independently |
| Changed files  | 5                | Number of files modified in this PR                            |
| Lines changed  | 219              | Total additions + deletions                                    |
| Owners         | @payments-team   | CODEOWNERS entries responsible for review                      |

2.2 Risk-gated review: ai-gate-review

The RADAR-style funnel where AI reviews every PR, computes a risk score, and produces a report.

# ai-gate-review.yaml — read-only pipeline
# review ? gate-check ? done (no merge, no approve, no PR comments)
- task_type: gate-check
  script:
    - |
      python3 -c "
      risk = json.load(open('/workspace/risk_score.json'))
      review = json.load(open('/workspace/review_result.json'))
      needs_approval = risk['score'] >= threshold or has_critical_findings
      # Writes gate_result.json — read-only, no PR changes
      print(f'Gate: {\"needs-approval\" if needs_approval else \"safe-to-merge\"}')"

Low-risk PRs are flagged “safe-to-merge” in the report. High-risk ones are flagged “needs-approval” with the specific reason. The risk score itself is a RADAR-style weighted composite across six dimensions (scripts/mq/risk_score.py).

2.3 The merge queue core

As a Formicary DAG it uses real fan-out with fork_job_type where each scope lane runs as its own child job:

# ai-merge-queue.yaml (excerpt)
job_type: ai-merge-queue
cron_trigger: "0 2 * * *"   # once daily at 2am; bump frequency when queue fills up
max_concurrency: 1

tasks:
- task_type: collect
  environment:
    TARGET_BRANCH: "{{.TargetBranch}}"   # filter to PRs targeting this branch
  script:
    - python -m scripts.mq.collect_ready   # reads TARGET_BRANCH from env

- task_type: group
  script:
    - python -m scripts.mq.group_by_scope   # produces hierarchical risk-tier lanes

- task_type: analyze
  script:
    - python -m scripts.mq.analyze --skill ygs-merge-queue

- task_type: report
  script:
    - mkdir -p /workspace/reports
    - python -m scripts.mq.report

Invoke from Slack with a target branch:

@bot mq --target stage          # short alias
@bot mq --target prod --repo org/my-repo

Each ai-mq-lane child job runs independently:

2.4 Work type distribution

The simulation discussed earlier has two knobs: defect rate and pipeline duration. The MQ report measures this from an actual PR data. For example, work type classification breaks every open PR into one of eight types: feature, bug, security, refactor, chore, test, docs, or unknown. It then measures defect rate, bug ratio, chore+refactor fraction.

What the report shows:

### Work Type Distribution
| Type      | Count | %     | Signal                      |
|-----------|-------|-------|-----------------------------|
| ? feature | 45    | 18.0% | new functionality           |
| ? bug     | 22    | 8.8%  | defect indicator            |
| ? security| 5     | 2.0%  | defect indicator (security) |
| ?? refactor | 30    | 12.0% | tech debt reduction         |
| ? chore   | 15    | 6.0%  | maintenance / KTLO          |
| ? test    | 120   | 48.0% | quality investment          |
| ? docs    | 3     | 1.2%  | documentation               |
| ? unknown | 10    | 4.0%  | unclassified                |

> ? Defect rate proxy: 10.8% (27 bug+security PRs / 250 total) — ~1-in-9.
> At batch size ~10, est. batch success ? 31% (moderate defect rate).
> Feature:Bug ratio = 1.7:1 — below 3:1, team spending significant effort on defect repair.

3. Quality at the Source

The defect-rate knob from earlier simulation is the hard to move but it is a high impact knob. Teams that have adopted agentic engineering at scale share a common pattern: an 8-phase SDLC that wraps AI capabilities with human judgment.

3.1 Structured specs reduce defect rate at source

Every ticket entering a sprint needs five things before an agent touches it:

  1. Why: one sentence on the customer problem
  2. Scope boundaries
  3. Aacceptance criteria in given/when/then form
  4. Definition of Done
  5. Environment matrix

An AI skill generates the structured ACs; a human validates and edits.

3.2 Design docs

Not every change needs a design doc. The decision tree:

SignalAction
New architecture pattern, public API change, multi-team impact, >1 sprint, auth/securityWrite a design doc
Bug fix, tests/config only, simple refactor in one file, dependency bumpNo design doc needed

When required, the design doc must address: risks, rollback strategy, MVP scope, testing strategy, and observability.

3.3 Tiered review

Not all PRs deserve the same review depth. A risk-based triage:

SignalsTier
Small diff, tests or config only, simple refactor, pre-PR gate passed cleanlyTier 1 Light pass
Auth/tokens/permissions, new or changed REST endpointTier 2 Full review

3.4 Contract testing + fuzz testing

The DORA/LeadDev findings showed that the contract testing before adopting AI reduces change failure rate. Contract testing catches the “clean-looking PR, silent cross-system break” failure mode.

I have another open source project api-mock-service, that can be used for contract and fuzz testing, e.g.,

# 1. Load an OpenAPI spec (or record live traffic through the proxy)
curl -X POST http://localhost:8080/_oapi -F "file=@openapi.yaml" -F "group=billing"

# 2. Run producer contract tests
curl -X POST "http://localhost:8080/_contracts/billing?baseUrl=http://billing:8080&executionTimes=3"

# 3. Run mutation tests (11 strategies)
curl -X POST "http://localhost:8080/_contracts/mutations/billing?baseUrl=http://billing:8080&executionTimes=5"

# 4. Export JUnit XML for CI
curl "http://localhost:8080/_contracts/billing/junit" > contract-results.junit.xml

3.5 Mutation strategies

Per-field mutations test individual field validation:

StrategyWhat it tests
Missing fieldsRequired field validation
Boundary valuesMin/max, empty strings, zero, MAX_INT
Malformed dataWrong types, invalid formats
Null fieldsNull handling, NPE prevention
Combinatorial nullsMulti-field invalid combinations
Format-specific boundariesDate edges, URL length, email format
Security injectionsSQLi, XSS, path traversal, SSTI, cmd injection, NoSQLi

Sequence-level mutations test stateful interactions:

StrategyWhat it tests
Request reorderingState machine correctness
Request duplicationIdempotency
Request omissionRequired step enforcement
Timing variationsRace conditions, timeouts

3.6 Contract testing as a Formicary job

The Formicary pipeline runs as a three-task Formicary DAG (docs/examples/ai-contract-test.yaml). Each task runs in a separate Kubernetes pod with api-mock-service as a sidecar:

job_type: ai-contract-test
description: "Contract validation + security fuzzing via api-mock-service proxy"
max_concurrency: 5
timeout: 3600s

variables:
  PRNumber:
    type: STRING
    required: false
  ServiceURL:
    type: STRING
    required: false
  Service:
    type: STRING
    required: false

tasks:
# --- RECORD ------------------------------------------------------------------
- task_type: record
  method: KUBERNETES
  timeout: 15m
  report_stdout: true
  host_network: true
  working_dir: /workspace
{{if .Service}}
  services:
    - name: "{{default "service-under-test" .ServiceName}}"
      alias: "{{default "service-under-test" .ServiceName}}"
      image: "{{.Service}}"
      memory_limit: "{{default "4G" .ServiceMemoryLimit}}"
      cpu_request: "{{default "250m" .ServiceCpuRequest}}"
    - name: api-mock-service
      alias: api-mock-service
      image: plexobject/api-mock-service:latest
      ports:
        - number: 8081
      memory_limit: "512Mi"
      cpu_request: "100m"
      command: ["/api-mock-service", "--httpPort", "8081", "--proxyPort", "8082", "--dataDir", "/workspace/recordings"]
      volumes:
        empty_dir:
          - name: workspace
            mount_path: /workspace
{{end}}
  container:
    image: plexobject/ai-dev-tools:latest
    image_pull_policy: Always
    cpu_request: "250m"
    memory_limit: 512Mi
    memory_request: 256Mi
    volumes:
      empty_dir:
        - name: workspace
          mount_path: /workspace
    env_from:
      - secret_ref: ai-dev-credentials
  environment:
    WORKSPACE_DIR: /workspace
    SERVICE_URL: "{{.ServiceURL}}"
    SERVICE_PORT: "{{default "8080" .ServicePort}}"
    MOCK_SERVICE_PORT: "{{default "8081" .MockServicePort}}"
    PROXY_PORT: "{{default "8082" .ProxyPort}}"
    AI_DEV_TOOLS_DEBUG: "{{.AiDevToolsDebug}}"
  script:
    - python -c "from scripts.common.bootstrap import ensure_debug_mode; ensure_debug_mode()" 2>/dev/null || true
    - touch /tmp/.adt_bootstrap_done
    - python -m scripts.contract.record
  artifacts:
    paths:
      - ./record_result.json
      - ./recordings
    expire_after: 24h
  on_completed: fuzz
  on_failed: notify-error

# --- FUZZ --------------------------------------------------------------------
- task_type: fuzz
  method: KUBERNETES
  timeout: 20m
  report_stdout: true
  host_network: true
  working_dir: /workspace
{{if .Service}}
  services:
    - name: "{{default "service-under-test" .ServiceName}}"
      alias: "{{default "service-under-test" .ServiceName}}"
      image: "{{.Service}}"
      memory_limit: "{{default "4G" .ServiceMemoryLimit}}"
      cpu_request: "{{default "250m" .ServiceCpuRequest}}"
    - name: api-mock-service
      alias: api-mock-service
      image: plexobject/api-mock-service:latest
      ports:
        - number: 8081
      memory_limit: "512Mi"
      cpu_request: "100m"
      command: ["/api-mock-service", "--httpPort", "8081", "--proxyPort", "8082", "--dataDir", "/workspace/recordings"]
      volumes:
        empty_dir:
          - name: workspace
            mount_path: /workspace
{{end}}
  container:
    image: plexobject/ai-dev-tools:latest
    image_pull_policy: Always
    cpu_request: "500m"
    memory_limit: 2G
    memory_request: 512Mi
    volumes:
      empty_dir:
        - name: workspace
          mount_path: /workspace
    env_from:
      - secret_ref: ai-dev-credentials
  dependencies:
    - record
  environment:
    WORKSPACE_DIR: /workspace
    SERVICE_URL: "{{.ServiceURL}}"
    SERVICE_PORT: "{{default "8080" .ServicePort}}"
    MOCK_SERVICE_PORT: "{{default "8081" .MockServicePort}}"
    PR_NUMBER: "{{.PRNumber}}"
    AI_DEV_TOOLS_DEBUG: "{{.AiDevToolsDebug}}"
  script:
    - python -c "from scripts.common.bootstrap import ensure_debug_mode; ensure_debug_mode()" 2>/dev/null || true
    - touch /tmp/.adt_bootstrap_done
    - python -m scripts.contract.fuzz
  artifacts:
    paths:
      - ./fuzz_result.json
      - ./fuzz_results.xml
      - ./contract_test_summary.json
    expire_after: 24h
  on_completed: report
  on_failed: report

The recording proxy is a sidecar container (plexobject/api-mock-service:latest) that shares the pod network namespace. Traffic flows:

test script ? localhost:8081 (proxy) ? localhost:8080 (service under test)
                                ?
                    api_contracts/**/*.yaml   (saved per HTTP interaction)

3.7 Skills

Two complementary skills in the you-got-skills library:

  • /ygs-contract-test: Detects API surface, sets up api-mock-service, runs contract validation and mutation testing.
  • /ygs-fuzz-test: Generates fuzz corpus from API specs, applies all mutation strategies plus 8 CWE-classified injection classes.

Here’s the fuzz skill’s decision table for endpoint discovery:

Signal foundAction
OpenAPI/Swagger specParse endpoints + schemas directly
Express routes (app.get/post)Extract route patterns from code
Flask decorators (@app.route)Extract route patterns from code
Go http.HandleFuncExtract route patterns from code
Spring @RequestMappingExtract route patterns from code
None of the aboveBLOCKED — no endpoints to fuzz

The 8 CWE-classified injection classes the fuzz skill exercises:

Injection classCWEExample payload
SQL injectionCWE-89' OR 1=1 --
XSSCWE-79<script>alert(1)</script>
Path traversalCWE-22../../etc/passwd
SSTICWE-1336{{7*7}}
Command injectionCWE-78; cat /etc/passwd
NoSQL injectionCWE-943{"$gt": ""}
LDAP injectionCWE-90`)(uid=))(
XXECWE-611<!DOCTYPE foo [<!ENTITY xxe SYSTEM "file:///etc/passwd">]>

3.8 Live example: OWASP WrongSecrets

OWASP WrongSecrets is a Java Spring Boot application intentionally seeded with secrets management vulnerabilities.

Trigger from Slack:

@sb-slack contract-test https://github.com/OWASP/wrongsecrets \
  --service jeroenwillemsen/wrongsecrets:latest-no-vault

The job produces output such as:

[contract] api-mock-service ready at http://localhost:8081/_health (HTTP 404)
[fuzz] contracts: succeeded=0 failed=0
[fuzz] recordings_dir=/workspace/recordings exists=True contracts_dir_exists=True
      yaml_count=21 first_3=[
        '/workspace/recordings/api_contracts/GET/Recordedroot--200-c1ec01cf.yaml',
        '/workspace/recordings/api_contracts/status/GET/...',
        '/workspace/recordings/api_contracts/debug/GET/...'
      ]
[fuzz] discovered 21 endpoints: [
  ('GET', '/status'), ('GET', '/debug'), ('GET', '/metrics'),
  ('GET', '/openapi.json'), ('GET', '/docs'), ('GET', '/'),
  ('GET', '/swagger'), ('GET', '/challenge/21'), ('GET', '/challenge/28'),
  ('GET', '/challenge/1'), ('GET', '/health'), ('GET', '/api'),
  ('GET', '/api/challenges'), ('GET', '/api/Challenges'), ('GET', '/admin'),
  ('GET', '/v1'), ('GET', '/actuator/env'), ('GET', '/actuator/info'),
  ('GET', '/actuator/beans'), ('GET', '/actuator/health'),
  ('GET', '/actuator/mappings')
]
[contract] probe sqli:/ ? 200 finding=False
[contract] probe path_trav:/ ? 404 finding=False
[contract] probe sqli:/status ? 404 finding=False
[contract] probe path_trav:/status ? 404 finding=False
... (40 probes total) ...
[fuzz] 21 endpoints, 0 findings, 0 critical
[fuzz] complete: iterations=40 findings=0 critical=0 status=PASS ams_used=True

3.9 What a skill looks like: /ygs-contract-test

Skills are structured Markdown files in the you-got-skills library. Here’s the decision table from /ygs-contract-test:

```yaml
# SKILL.md frontmatter
---
name: ygs-contract-test
description: API contract testing — record, derive contracts, detect breaking changes
argument-hint: "<service-or-pr> [--mode record|validate|diff]"
---
```

```markdown
## Step 1: Determine contract testing mode

| Mode     | When to use                          | What happens                        |
|----------|--------------------------------------|-------------------------------------|
| record   | First time, or updating baseline     | Run tests through proxy, capture    |
| validate | PR review, CI gate                   | Compare against existing contracts  |
| diff     | Breaking change detection            | Compare two contract versions       |

Default to `validate` if a baseline exists, otherwise `record`.

## Step 2: Record API interactions
  api-mock-service --mode record --proxy-port 8081 ...
  HTTP_PROXY=http://localhost:8081 pytest tests/integration/

## Step 3: Derive and validate contracts
  curl -X POST "http://localhost:8080/_contracts/{group}?baseUrl=..."

## Step 4: Run mutation testing (11 strategies)
  curl -X POST "http://localhost:8080/_contracts/mutations/{group}..."

## Step 5: Export and report
  Report DONE if catch rate > 80%. BLOCKED if contracts fail.
  Suggest: /ygs-fuzz-test for deeper coverage, /ygs-security-review for findings.

4. Speed at the Pipeline

The second knob from earlier simulation is pipeline speed, which is a cheaper lever. This section covers formal verification of the queue protocol itself, build/test optimization with real numbers.

4.1 Formal verification

Tests check specific inputs but formal verification proves properties hold over all possible inputs. For a concurrent system like a merge queue, the state space is too large to test exhaustively.

TLA+ specification of the merge queue

The spec lives at docs/examples/specs/merge_queue.tla and it models the core merge queue as a state machine:

Safety for Scope isolation

ScopeIsolation ==
    \A s1, s2 \in Scopes :
        s1 /= s2 => lanes[s1] \cap lanes[s2] = {}

Safety for Test before merge:

TestBeforeMerge ==
    \A pr \in PRs :
        prState[pr] = "merged" => testResults[pr] = "pass"

Liveness to ensure every PR eventually merges or is ejected:

Progress == \A pr \in PRs :
    prState[pr] = "queued" ~> (prState[pr] = "merged" \/ prState[pr] = "ejected")

Dafny verification of merge invariants

The Dafny spec at docs/examples/specs/verified_merge.dfy proves four properties at compile time: batch merging preserves main branch health, bisection correctly partitions PRs, risk scoring is monotonic, and shard partitioning loses no tests:

predicate MainBranchHealthy(mergedPRs: set<PR>)
{
  forall pr :: pr in mergedPRs ==> pr.testResult == Pass
}

lemma MergeBatchPreservesHealth(mainBranch: set<PR>, batch: set<PR>)
  requires MainBranchHealthy(mainBranch)
  requires forall pr :: pr in batch ==> pr.testResult == Pass
  ensures MainBranchHealthy(mainBranch + batch)
{
  // Proof is automatic: union of two sets where all elements
  // satisfy the predicate still satisfies the predicate.
}

4.2 Test-impact analysis: scripts/mq/test_impact.py

The default behaviour is deliberate: always run the full test suite, partitioned into balanced shards. Diff-scoped runs are opt-in via --diff-scope and they require an explicit --head <branch> to define what to diff against.

# Default: full suite, 8 shards, no diff analysis
python -m scripts.mq.test_impact --pr-number main --num-shards 8

# Explicit diff-scope: numeric PR (base branch auto-extracted from GitHub API)
python -m scripts.mq.test_impact --pr-number 42 --num-shards 8 --diff-scope

# Explicit diff-scope: branch name (must supply --head as base to compare against)
python -m scripts.mq.test_impact --pr-number feature/my-branch \
    --head main --num-shards 8 --diff-scope

The script performs following operation:

  1. Fetches changed files via gh pr view --json
  2. Detects language from file extensions (Python, Go, TypeScript, Java, Kotlin, Ruby, C#, Rust)
  3. Maps to test files using naming conventions
  4. Traverses imports one level deep to find transitive dependents
  5. Partitions into balanced shards using greedy bin-packing with historical timing data
  6. Falls back to the full suite if no tests map to changed files

4.3 Parallel fan-out: ai-parallel-test

Test-impact analysis feeds directly into Formicary’s fan-out for parallel shard execution (docs/examples/ai-parallel-test.yaml):

# analyze task: runs test_impact.py, emits ::add-job-context TestShards::
- task_type: analyze
  script:
    - python -m scripts.mq.clone_pr --pr-number {{.PRNumber}}
    - python -m scripts.mq.test_impact --pr-number {{.PRNumber}}
    # TestShards is now in job context — fan-out reads it directly

# Fan-out: one task per test shard, running in parallel
# 4 CPU cores + 16G memory per shard ? real intra-shard parallelism
- task_type: run-tests
  container:
    cpu_request: "4000m"
    memory_limit: 16G
  fan_out:
    source: TestShards
    item_var: shard
    max_parallel: {{.MaxShards}}  # unquoted integer — YAML strict types
    fail_fast: false
  script:
    - python -m scripts.mq.clone_pr --pr-number {{.PRNumber}}
    - python -m scripts.mq.run_scoped_ci --shard "{{.shard}}"

Each shard runs independently on its own Kubernetes pod. The polyglot test runner (scripts/mq/run_scoped_ci.py) detects the project type from marker files and builds the right command.

Here is a sample output:

## Test Impact Analysis
**191** / 191 tests selected (**0%** reduction) across **2** shards

## Test Results
? **87** / 87 passed
  ?? 130s wall clock across 2 shards

### Shard Performance
| Shard | Tests | Passed | Failed | Duration | Status |
|-------|-------|--------|--------|----------|--------|
| 1     | 44    | 44     | 0      | 130.5s   | ?     |
| 0     | 43    | 43     | 0      | 115.4s   | ?     |

> **Parallel speedup:** 246s sequential ? 130s parallel (1.9x across 2 shards)

### Slowest Tests
| Duration | Test                          |
|----------|-------------------------------|
| 30.01s   | `TestInitEnabled`             |
| 30.01s   | `TestInitEnabled`             |
| 10.02s   | `Test_ShouldAsyncWithSleep`   |
| 10.01s   | `Test_ShouldAsyncWithSleep`   |
|  4.68s   | `Test_EncryptDecrypt`         |
|  2.45s   | `Test_ShouldRealGet`          |
|  2.28s   | `TestSharedSubscription`      |

### Test Health Insights
- ? Pass rate: 100% (87 tests)
- ? Shard balance: 12% imbalance (115s – 130s)
- ? 4 slow tests likely using real sleeps — consider mocking time or
  reducing timeouts: `TestInitEnabled` (30.0s), `Test_ShouldAsyncWithSleep` (10.0s)
- ?? 3 slow test names appear in multiple shards (possible test duplication):
  `TestInitEnabled`, `Test_ShouldAsyncWithSleep`, `TestSharedSubscription`

4.4 Deployment risk

The merge queue optimizes everything up to the point where the commit lands on main but merged PRs still have to ship. The original simulation from Joe Magerramov’s formula models the probability that a merge batch is clean and continuous delivery. In my simulation I added support for batch PRs as it’s used in many organizations. But this also increases risk, e.g., a merge queue with a 95% batch success rate can still lead to a weekly release train that ships clean only 36% of the time. The formula extends naturally:

release_success = (1 - defect_rate) ^ prs_per_release

CD is the special case where prs_per_release = 1, e.g.,

Defect rateCD (1 PR)Daily (10 PRs)Weekly (50 PRs)Bi-weekly (100 PRs)
0.5%99.5%95.1%77.8%60.6%
1%99.0%90.4%60.5%36.6%
2%98.0%81.7%36.4%13.3%
5%95.0%59.9%7.7%0.6%

At 2% defect rate, a CD pipeline has a 98% clean deployment rate. The same team on a weekly train: 36%. Two-thirds of their releases ship with at least one defect.

The rollback trap

Suppose release R1 ships with a hidden bug. By the time monitoring detects it, R2 and R3 have already deployed. Rolling back to pre-R1 means reverting R2 and R3 too.

Stacked releasesRollback feasibilityRecovery strategyMTTR multiplier
1? StraightforwardRollback1.0x
2-3?? Costly but possibleRollback (with intermediate revert)1.5x
?4? ImpracticalRoll-forward (fix-forward only)2.5x

One bad PR poisons the entire train and the release sits in an unknown state while someone bisects 50 changes to find the culprit.

Deployment maturity

Following are best practices from the industry for improving deployment maturity:

DimensionWeightWhat it enables
Automated testing2.0xCI on every PR catches defects pre-merge
Canary deployment2.0xProgressive rollout catches prod-only failures
Automated rollback1.5xInstant revert on anomaly detection
Observability1.5xAlerting + metrics detect failures fast
Wave deployment1.0xStage ? preprod ? prod progression
Feature flags1.0xDecouple deploy from release
Blue/green0.5xZero-downtime deploy infrastructure

The maturity score translates into a risk multiplier — a low multiplier means your infrastructure catches most defects before they reach users:

adjusted_success = release_success + (1 - release_success) × (1 - risk_multiplier)

The MQ report now shows two critical markers:

  • ? Blue line: your current position given your defect rate, batch size, and deployment maturity.
  • ? Red line: the calamity threshold: the batch size where success drops below 70%.

The calamity threshold formula:

max_safe_batch = log(0.70) / log(1 - defect_rate)

The deployment risk model points to three levers:

  1. Reduce batch size. Move from weekly to daily trains.
  2. Invest in deployment maturity. Automated testing (weight 2.0x) and canary deployment (weight 2.0x) are the two highest-leverage dimensions.
  3. Prevent release stacking. Each additional stacked release multiplies recovery cost.

5. The Learning Loop

5.1 Multi-agent coordination

The merge queue is naturally an actor system where each PR is an actor with a lifecycle (queued ? batched ? testing ? merged/ejected). The properties that matter at scale:

  • Location transparency: a PR doesn’t care which CI runner tests it
  • Supervision: when a batch fails, the supervisor (bisect task) localizes the fault and recovers
  • Isolation: scope-aware lanes are isolation boundaries

Better decomposition

Similar to the microservice architecture, the path forward is better decomposition:

  • Smaller PRs
  • Better module boundaries
  • Formal verification
  • Automated scope detection
  • Dependent PR grouping, e.g., group related PRs so if one fails, the entire group is ejected.

5.2 The skills improvement flywheel

5.3 Metrics

The four DORA keys (deployment frequency, lead time, change failure rate, failed deployment recovery time) are table stakes. At agent-scale, you need additional metrics like:

  • PR size
  • PR pickup time Stale PRs conflict more often and carry higher bisection cost in the merge queue.
  • AI PR acceptance rate.
  • Rework rate.
  • Cycle time decomposition break first-commit-to-deploy into pickup, review, merge, deploy sub-phases.

5.4 Path forward

The bottom line from all the evidence in this post:

  1. Don’t adopt AI at agent-scale before your delivery system can handle agent-scale quality.
  2. Pipeline speed is the cheaper lever. Defect rate has diminishing returns; pipeline duration has an order of magnitude of headroom.
  3. The merge queue is not a queue. It’s a coordination problem: scope-aware routing, risk-gated review, speculative batching, bisection-based fault recovery, and a learning loop.
  4. Formal verification catches what tests can’t. For concurrent protocols, tests check paths and TLA+ checks invariants.
  5. Specs reduce defects at source. Structured acceptance criteria, design docs for high-impact changes, and tiered review based on blast radius.

6. Getting Started

Everything above is code in three repos: formicary (workflow engine), ai-dev-tools (Python scripts), and you-got-skills (Claude skill library). Here’s how to test each piece.

Prerequisites

# 1. Deploy Formicary to Kubernetes (single-container: queen + ant + storage)
kubectl apply -f k8s.yaml
kubectl port-forward svc/formicary 7777:7777 19000:19000

# 2. Get an API token from the UI: http://localhost:7777 ? Profile ? API Token
export FORMICARY_TOKEN="<jwt>"
export GH_TOKEN="<github-pat>"
export SLACK_CHANNEL="your-channel"

# 3. Deploy all workflows + set configs
cd docs/examples
./deploy-ai-workflows.sh --create-k8s-secret --set-configs \
  --gh-org bhatti --gh-repo todo-api-errors --bedrock

Test 1: Scope router

In Slack, mention the bot in your configured channel:

@bot scope main --repo https://github.com/org/repo

This prompt triggers ai-scope-router workflow and then creates a report showing the scope, blast radius, coverage, historical data for defects, security and other metrics.

Test 2: Parallel test with test-impact analysis

@bot parallel-test --repo https://github.com/org/repo main
@bot parallel-test --repo https://github.com/org/repo --branch develop

This prompt triggers ai-parallel-test workflow that fans out multiple shards for parallel tests and then creates a report showing the test results.

Test 3: Risk-gated review

@bot gate-review https://github.com/org/repo/pull/1

This prompt triggers ai-gate-review workflow and generates a report showing blast radius, security, test coverage and other metrics.

Test 4: Contract testing

@bot contract-test 1 --service myorg/my-api:latest

This prompt triggers ai-contract-test workflow and executes contract test against the target service. It then generates a report with the test results.

Test 5: Merge Queue Analysis

@bot mq https://github.com/org/repo

This prompt triggers ai-merge-queue workflow, evaluating health of the merge queue and calculating the risk/blast radius/complexity of all open PRs. Here is a sample snippet from the merge queue report:

? 250 open PRs across 22 scope lanes — repo: acme/platform

MetricValue
Total open PRs250
Scope lanes22
High-risk PRs140
Needs human review148
Conflict risk lanes1

?? 148 PR(s) flagged for human review before merge (high blast-radius, failed CI, or sensitive paths detected).

?? 1 lane(s) have conflict risk multiple PRs touching overlapping paths. Merge one at a time.

Overall status: ? Unstable (basis: 74.8% PRs aged >48h — CI status unavailable from API)

MetricValue
Total PRs in queue250
CI failure rateN/A — not reported by API (verify in CI dashboard)
PRs aged >48h187 (74.8%)
High blast-radius PRs140
Medium blast-radius PRs82
Approx batch size (PRs/lane)10.0

Unstable: 74.8% of PRs are aged >48h. Queue is filling faster than it drains — unblock stale PRs or shrink batch size.

CategoryPRsBug PRsHigh-BlastHotspot
? ui12211??? yes
? data323??? yes
? api315??? yes
? authn_authz315??? yes
? security153??? yes
sre121??—
backend30??—
config10??—
unknown30——

Test 6: PR Audit

@bot pr-audit

This prompt triggers ai-gh-pr-audit workflow, evaluating the risk/blast radius/complexity of all merged PRs. Here is a sample snippet from the pr audit report:

#PRIssueAuthorTitleCatTypeBlastRiskLOCFilesCx
1#48401PROJ-44822jsmithAuthCoordinator for REST OAuth? authn_authz? securityhigh? 721,84723? high
2#47228AI-4931aleefix(ai): allow regional Bedrock pr…? data? bughigh? 5889214? medium
3#48768DL-2875mchenImplement more fine-grained IAM roles? authn_authz? featurehigh? 4563411? medium
4#48909—jdoeSSRF guard for external MCP connectorui? bugmedium? 323128? low
5#48827PLAT-16031kpatel[Flaky Test] fix race in bottlenec…api?? refactorlow? 12873? low
…

Related Reading

No Comments

No comments yet.

RSS feed for comments on this post. TrackBack URL

Sorry, the comment form is closed at this time.

Powered by WordPress