Here’s how to define where AI fits, reduce cognitive and intent debt, and keep your delivery pipeline stable.
1. The AI Genie
AI coding agents can produce code at an incredible speed but without the quality, you will find that it quickly derails the release train. The existing CI and code-review processes were designed for human-paced changes but this doesn’t keep up when you are pushing code at 10x and you end up with the release problem.

Here is how some of the companies are trying to solve merge queue problems:
| Source | Finding |
|---|---|
| Linear | Cut PR wait time via faster runners, tsc -> tsgo, sparse checkout, additional test shards |
| Atlassian | 70+ large repos on merge queues, PR-level CI answers “does this work alone,” not “does this still work with everything” |
| Nirvana | Stateless speculative merge queue. |
| Mergify | Speculative testing, batching, and scope-aware parallel lanes |
| Autonoma | Parallel AI subagents multiply conflict surface area unless generation is batched and merges serialized |
Here is what I found how AI is affecting the quality:
| Source | Finding |
|---|---|
| DORA 2025 | PRs merged per developer +98%. Incidents per PR +243%. Bugs per developer +54% |
| Cortex 2026 | PRs per author +20% YoY. Incidents per PR +23.5%. Change failure rate ~30% |
| GitHub Octoverse 2024 | 92% of developers use AI coding tools. GitHub Actions CI/CD minutes up 169% |
| GitClear (211M lines) | Code churn doubled. Refactored code fell 24.1% ? 9.5%. Duplicated blocks rose 8x |
| Uplevel (800 devs, 3 months) | +41% bugs after Copilot adoption. No improvement in cycle time or throughput |
| Opsera (250k+ devs, 60+ orgs) | AI PRs sit 4.6x longer in review queues; duplication 10.5% ? 13.5% |
Joe Magerramov’s post shows how to build a small Monte Carlo simulator for modeling merge queues. Joe showed how CI/CD becomes a traffic jam, not just a queue. CI/CD batches are cumulative so one defect forces a revert and next batch is affected. I built my own simulator based on Joe’s model (see https://github.com/bhatti/simulators) with additional support for the batched PRs release. It models defect rate with the pipeline duration at 100 commits per day:

The key lesson is that you need to either lower the defect rate per commit or shorten pipeline duration. This is not an easy task, I have encountered a large pipeline duration at many organizations due to large mono-repos, a large codebase with millions of LOC, and large test suite with a long vaidation cycle. The AI agents makes it worse with parallel changes that might conflict resulting in cognitive and intent debt for engineers because no one can track all changes. I explained some of these concepts in my earlier blog and showed how learning feedback loops can be used to build resilient software factories. In order to build end to end agentic SDLC process, you need to define what is AI responsible for and what are the roles for humans.
| Quadrant | Role |
|---|---|
| Human Real-time | Decisions that require judgment, context, accountability |
| Human Async | Review that needs thought but not immediacy |
| AI Real-time | Assistance that augments human work in the moment |
| AI Background | Autonomous work that runs without blocking humans |
2. The Architecture
I built an orchestration engine Formicary to create data pipelines and CI/CD processes many years ago but I have been using it for driving AI driven workflows. It defines simple primitives to build DAG tasks with exit-code routing, artifact handoff, and fan-out. Here are a few approaches that I am using with the AI driven workflows:
- Speculative merge queue (test PR N as if N-1 already merged)
- Scope-aware parallel lanes
- Incremental builds, test-impact analysis, sparse checkout
- Deterministic gates a green build can’t waive
- Independent AI review
- Contract testing + canary
- Formal verification of queue invariants (TLA+, Dafny)
- API fuzz testing
- Learning flywheel: merge -> extract learnings -> next run reads them -> audit proposes skill changes
Here is how I use Formicary with a CLI toolkit and skillsThe system is split across three repos, each with a clear responsibility boundary:

Here are the core design principles
- Declarative DAG with a task block that uses
on_exit_codefor routing - Harness + sandbox + skills as separate layers
- File-based state handoff
- The learning flywheel closes the loop on review time and merge-time

2.1 Scope router: ai-scope-router
Runs right after create-pr in the existing pipeline and computes a scope key from touched paths. The backing script (scripts/mq/scope_router.py) computes blast radius from line counts and module count, and labels the PR.
# ai-scope-router.yaml (excerpt — full file in docs/examples/)
job_type: ai-scope-router
max_concurrency: 20
timeout: 600s
tasks:
- task_type: classify
method: KUBERNETES
script:
- python -m scripts.mq.scope_router --pr-number {{.PRNumber}}
- python -m scripts.mq.risk_score --pr-number {{.PRNumber}}
on_completed: route
- task_type: route
script:
- |
python3 -c "
import json, os
scope = json.load(open('/workspace/scope.json'))
risk = json.load(open('/workspace/risk_score.json'))
decision = {
'scope_key': scope['scope'],
'risk_tier': risk['tier'],
'lane': scope['scope'],
'requires_approval': risk.get('requires_human_approval', False)
}
json.dump(decision, open('/workspace/route_decision.json', 'w'), indent=2)
"
on_completed: done
Here’s what the Slack report looks like for a PR:
# PR Review Report — PR #4091
## Review Findings
? No issues found
## Risk Score
? MEDIUM (score 25.5/100) — Standard review needed; test independently before merge
| Dimension | Score | Weight | Evidence |
|-----------------|-------|--------|---------------------------------------------------|
| Size | 7/10 | 1.5× | 219 lines (+147/?72) |
| File Count | 2/10 | 1.0× | 5 files changed |
| Blast Radius | 5/10 | 2.0× | blast=medium, scope=payments-service |
| Sensitive Paths | 0/10 | 2.5× | no sensitive files detected |
| Test Coverage | 0/10 | 1.5× | good coverage (test:source ?1:1) |
| Historical | 3/10 | 1.0× | ?? no defect history available — neutral default |
## Scope
| Field | Value | Description |
|----------------|------------------|----------------------------------------------------------------|
| Scope | payments-service | All changes owned by payments-service — can merge in dedicated lane |
| Blast radius | medium | Moderate change (51–300 lines or 2 modules) — test independently |
| Changed files | 5 | Number of files modified in this PR |
| Lines changed | 219 | Total additions + deletions |
| Owners | @payments-team | CODEOWNERS entries responsible for review |
2.2 Risk-gated review: ai-gate-review
The RADAR-style funnel where AI reviews every PR, computes a risk score, and produces a report.
# ai-gate-review.yaml — read-only pipeline
# review ? gate-check ? done (no merge, no approve, no PR comments)
- task_type: gate-check
script:
- |
python3 -c "
risk = json.load(open('/workspace/risk_score.json'))
review = json.load(open('/workspace/review_result.json'))
needs_approval = risk['score'] >= threshold or has_critical_findings
# Writes gate_result.json — read-only, no PR changes
print(f'Gate: {\"needs-approval\" if needs_approval else \"safe-to-merge\"}')"
Low-risk PRs are flagged “safe-to-merge” in the report. High-risk ones are flagged “needs-approval” with the specific reason. The risk score itself is a RADAR-style weighted composite across six dimensions (scripts/mq/risk_score.py).
2.3 The merge queue core
As a Formicary DAG it uses real fan-out with fork_job_type where each scope lane runs as its own child job:
# ai-merge-queue.yaml (excerpt)
job_type: ai-merge-queue
cron_trigger: "0 2 * * *" # once daily at 2am; bump frequency when queue fills up
max_concurrency: 1
tasks:
- task_type: collect
environment:
TARGET_BRANCH: "{{.TargetBranch}}" # filter to PRs targeting this branch
script:
- python -m scripts.mq.collect_ready # reads TARGET_BRANCH from env
- task_type: group
script:
- python -m scripts.mq.group_by_scope # produces hierarchical risk-tier lanes
- task_type: analyze
script:
- python -m scripts.mq.analyze --skill ygs-merge-queue
- task_type: report
script:
- mkdir -p /workspace/reports
- python -m scripts.mq.report
Invoke from Slack with a target branch:
@bot mq --target stage # short alias
@bot mq --target prod --repo org/my-repo
Each ai-mq-lane child job runs independently:

2.4 Work type distribution
The simulation discussed earlier has two knobs: defect rate and pipeline duration. The MQ report measures this from an actual PR data. For example, work type classification breaks every open PR into one of eight types: feature, bug, security, refactor, chore, test, docs, or unknown. It then measures defect rate, bug ratio, chore+refactor fraction.
What the report shows:
### Work Type Distribution
| Type | Count | % | Signal |
|-----------|-------|-------|-----------------------------|
| ? feature | 45 | 18.0% | new functionality |
| ? bug | 22 | 8.8% | defect indicator |
| ? security| 5 | 2.0% | defect indicator (security) |
| ?? refactor | 30 | 12.0% | tech debt reduction |
| ? chore | 15 | 6.0% | maintenance / KTLO |
| ? test | 120 | 48.0% | quality investment |
| ? docs | 3 | 1.2% | documentation |
| ? unknown | 10 | 4.0% | unclassified |
> ? Defect rate proxy: 10.8% (27 bug+security PRs / 250 total) — ~1-in-9.
> At batch size ~10, est. batch success ? 31% (moderate defect rate).
> Feature:Bug ratio = 1.7:1 — below 3:1, team spending significant effort on defect repair.
3. Quality at the Source
The defect-rate knob from earlier simulation is the hard to move but it is a high impact knob. Teams that have adopted agentic engineering at scale share a common pattern: an 8-phase SDLC that wraps AI capabilities with human judgment.

3.1 Structured specs reduce defect rate at source
Every ticket entering a sprint needs five things before an agent touches it:
- Why: one sentence on the customer problem
- Scope boundaries
- Aacceptance criteria in given/when/then form
- Definition of Done
- Environment matrix
An AI skill generates the structured ACs; a human validates and edits.
3.2 Design docs
Not every change needs a design doc. The decision tree:
| Signal | Action |
|---|---|
| New architecture pattern, public API change, multi-team impact, >1 sprint, auth/security | Write a design doc |
| Bug fix, tests/config only, simple refactor in one file, dependency bump | No design doc needed |
When required, the design doc must address: risks, rollback strategy, MVP scope, testing strategy, and observability.
3.3 Tiered review
Not all PRs deserve the same review depth. A risk-based triage:
| Signals | Tier |
|---|---|
| Small diff, tests or config only, simple refactor, pre-PR gate passed cleanly | Tier 1 Light pass |
| Auth/tokens/permissions, new or changed REST endpoint | Tier 2 Full review |
3.4 Contract testing + fuzz testing
The DORA/LeadDev findings showed that the contract testing before adopting AI reduces change failure rate. Contract testing catches the “clean-looking PR, silent cross-system break” failure mode.

I have another open source project api-mock-service, that can be used for contract and fuzz testing, e.g.,
# 1. Load an OpenAPI spec (or record live traffic through the proxy)
curl -X POST http://localhost:8080/_oapi -F "file=@openapi.yaml" -F "group=billing"
# 2. Run producer contract tests
curl -X POST "http://localhost:8080/_contracts/billing?baseUrl=http://billing:8080&executionTimes=3"
# 3. Run mutation tests (11 strategies)
curl -X POST "http://localhost:8080/_contracts/mutations/billing?baseUrl=http://billing:8080&executionTimes=5"
# 4. Export JUnit XML for CI
curl "http://localhost:8080/_contracts/billing/junit" > contract-results.junit.xml
3.5 Mutation strategies
Per-field mutations test individual field validation:
| Strategy | What it tests |
|---|---|
| Missing fields | Required field validation |
| Boundary values | Min/max, empty strings, zero, MAX_INT |
| Malformed data | Wrong types, invalid formats |
| Null fields | Null handling, NPE prevention |
| Combinatorial nulls | Multi-field invalid combinations |
| Format-specific boundaries | Date edges, URL length, email format |
| Security injections | SQLi, XSS, path traversal, SSTI, cmd injection, NoSQLi |
Sequence-level mutations test stateful interactions:
| Strategy | What it tests |
|---|---|
| Request reordering | State machine correctness |
| Request duplication | Idempotency |
| Request omission | Required step enforcement |
| Timing variations | Race conditions, timeouts |
3.6 Contract testing as a Formicary job
The Formicary pipeline runs as a three-task Formicary DAG (docs/examples/ai-contract-test.yaml). Each task runs in a separate Kubernetes pod with api-mock-service as a sidecar:
job_type: ai-contract-test
description: "Contract validation + security fuzzing via api-mock-service proxy"
max_concurrency: 5
timeout: 3600s
variables:
PRNumber:
type: STRING
required: false
ServiceURL:
type: STRING
required: false
Service:
type: STRING
required: false
tasks:
# --- RECORD ------------------------------------------------------------------
- task_type: record
method: KUBERNETES
timeout: 15m
report_stdout: true
host_network: true
working_dir: /workspace
{{if .Service}}
services:
- name: "{{default "service-under-test" .ServiceName}}"
alias: "{{default "service-under-test" .ServiceName}}"
image: "{{.Service}}"
memory_limit: "{{default "4G" .ServiceMemoryLimit}}"
cpu_request: "{{default "250m" .ServiceCpuRequest}}"
- name: api-mock-service
alias: api-mock-service
image: plexobject/api-mock-service:latest
ports:
- number: 8081
memory_limit: "512Mi"
cpu_request: "100m"
command: ["/api-mock-service", "--httpPort", "8081", "--proxyPort", "8082", "--dataDir", "/workspace/recordings"]
volumes:
empty_dir:
- name: workspace
mount_path: /workspace
{{end}}
container:
image: plexobject/ai-dev-tools:latest
image_pull_policy: Always
cpu_request: "250m"
memory_limit: 512Mi
memory_request: 256Mi
volumes:
empty_dir:
- name: workspace
mount_path: /workspace
env_from:
- secret_ref: ai-dev-credentials
environment:
WORKSPACE_DIR: /workspace
SERVICE_URL: "{{.ServiceURL}}"
SERVICE_PORT: "{{default "8080" .ServicePort}}"
MOCK_SERVICE_PORT: "{{default "8081" .MockServicePort}}"
PROXY_PORT: "{{default "8082" .ProxyPort}}"
AI_DEV_TOOLS_DEBUG: "{{.AiDevToolsDebug}}"
script:
- python -c "from scripts.common.bootstrap import ensure_debug_mode; ensure_debug_mode()" 2>/dev/null || true
- touch /tmp/.adt_bootstrap_done
- python -m scripts.contract.record
artifacts:
paths:
- ./record_result.json
- ./recordings
expire_after: 24h
on_completed: fuzz
on_failed: notify-error
# --- FUZZ --------------------------------------------------------------------
- task_type: fuzz
method: KUBERNETES
timeout: 20m
report_stdout: true
host_network: true
working_dir: /workspace
{{if .Service}}
services:
- name: "{{default "service-under-test" .ServiceName}}"
alias: "{{default "service-under-test" .ServiceName}}"
image: "{{.Service}}"
memory_limit: "{{default "4G" .ServiceMemoryLimit}}"
cpu_request: "{{default "250m" .ServiceCpuRequest}}"
- name: api-mock-service
alias: api-mock-service
image: plexobject/api-mock-service:latest
ports:
- number: 8081
memory_limit: "512Mi"
cpu_request: "100m"
command: ["/api-mock-service", "--httpPort", "8081", "--proxyPort", "8082", "--dataDir", "/workspace/recordings"]
volumes:
empty_dir:
- name: workspace
mount_path: /workspace
{{end}}
container:
image: plexobject/ai-dev-tools:latest
image_pull_policy: Always
cpu_request: "500m"
memory_limit: 2G
memory_request: 512Mi
volumes:
empty_dir:
- name: workspace
mount_path: /workspace
env_from:
- secret_ref: ai-dev-credentials
dependencies:
- record
environment:
WORKSPACE_DIR: /workspace
SERVICE_URL: "{{.ServiceURL}}"
SERVICE_PORT: "{{default "8080" .ServicePort}}"
MOCK_SERVICE_PORT: "{{default "8081" .MockServicePort}}"
PR_NUMBER: "{{.PRNumber}}"
AI_DEV_TOOLS_DEBUG: "{{.AiDevToolsDebug}}"
script:
- python -c "from scripts.common.bootstrap import ensure_debug_mode; ensure_debug_mode()" 2>/dev/null || true
- touch /tmp/.adt_bootstrap_done
- python -m scripts.contract.fuzz
artifacts:
paths:
- ./fuzz_result.json
- ./fuzz_results.xml
- ./contract_test_summary.json
expire_after: 24h
on_completed: report
on_failed: report
The recording proxy is a sidecar container (plexobject/api-mock-service:latest) that shares the pod network namespace. Traffic flows:
test script ? localhost:8081 (proxy) ? localhost:8080 (service under test)
?
api_contracts/**/*.yaml (saved per HTTP interaction)
3.7 Skills
Two complementary skills in the you-got-skills library:
/ygs-contract-test: Detects API surface, sets up api-mock-service, runs contract validation and mutation testing./ygs-fuzz-test: Generates fuzz corpus from API specs, applies all mutation strategies plus 8 CWE-classified injection classes.
Here’s the fuzz skill’s decision table for endpoint discovery:
| Signal found | Action |
|---|---|
| OpenAPI/Swagger spec | Parse endpoints + schemas directly |
| Express routes (app.get/post) | Extract route patterns from code |
| Flask decorators (@app.route) | Extract route patterns from code |
| Go http.HandleFunc | Extract route patterns from code |
| Spring @RequestMapping | Extract route patterns from code |
| None of the above | BLOCKED — no endpoints to fuzz |
The 8 CWE-classified injection classes the fuzz skill exercises:
| Injection class | CWE | Example payload |
|---|---|---|
| SQL injection | CWE-89 | ' OR 1=1 -- |
| XSS | CWE-79 | <script>alert(1)</script> |
| Path traversal | CWE-22 | ../../etc/passwd |
| SSTI | CWE-1336 | {{7*7}} |
| Command injection | CWE-78 | ; cat /etc/passwd |
| NoSQL injection | CWE-943 | {"$gt": ""} |
| LDAP injection | CWE-90 | `)(uid=))( |
| XXE | CWE-611 | <!DOCTYPE foo [<!ENTITY xxe SYSTEM "file:///etc/passwd">]> |
3.8 Live example: OWASP WrongSecrets
OWASP WrongSecrets is a Java Spring Boot application intentionally seeded with secrets management vulnerabilities.
Trigger from Slack:
@sb-slack contract-test https://github.com/OWASP/wrongsecrets \
--service jeroenwillemsen/wrongsecrets:latest-no-vault
The job produces output such as:
[contract] api-mock-service ready at http://localhost:8081/_health (HTTP 404)
[fuzz] contracts: succeeded=0 failed=0
[fuzz] recordings_dir=/workspace/recordings exists=True contracts_dir_exists=True
yaml_count=21 first_3=[
'/workspace/recordings/api_contracts/GET/Recordedroot--200-c1ec01cf.yaml',
'/workspace/recordings/api_contracts/status/GET/...',
'/workspace/recordings/api_contracts/debug/GET/...'
]
[fuzz] discovered 21 endpoints: [
('GET', '/status'), ('GET', '/debug'), ('GET', '/metrics'),
('GET', '/openapi.json'), ('GET', '/docs'), ('GET', '/'),
('GET', '/swagger'), ('GET', '/challenge/21'), ('GET', '/challenge/28'),
('GET', '/challenge/1'), ('GET', '/health'), ('GET', '/api'),
('GET', '/api/challenges'), ('GET', '/api/Challenges'), ('GET', '/admin'),
('GET', '/v1'), ('GET', '/actuator/env'), ('GET', '/actuator/info'),
('GET', '/actuator/beans'), ('GET', '/actuator/health'),
('GET', '/actuator/mappings')
]
[contract] probe sqli:/ ? 200 finding=False
[contract] probe path_trav:/ ? 404 finding=False
[contract] probe sqli:/status ? 404 finding=False
[contract] probe path_trav:/status ? 404 finding=False
... (40 probes total) ...
[fuzz] 21 endpoints, 0 findings, 0 critical
[fuzz] complete: iterations=40 findings=0 critical=0 status=PASS ams_used=True
3.9 What a skill looks like: /ygs-contract-test
Skills are structured Markdown files in the you-got-skills library. Here’s the decision table from /ygs-contract-test:
```yaml
# SKILL.md frontmatter
---
name: ygs-contract-test
description: API contract testing — record, derive contracts, detect breaking changes
argument-hint: "<service-or-pr> [--mode record|validate|diff]"
---
```
```markdown
## Step 1: Determine contract testing mode
| Mode | When to use | What happens |
|----------|--------------------------------------|-------------------------------------|
| record | First time, or updating baseline | Run tests through proxy, capture |
| validate | PR review, CI gate | Compare against existing contracts |
| diff | Breaking change detection | Compare two contract versions |
Default to `validate` if a baseline exists, otherwise `record`.
## Step 2: Record API interactions
api-mock-service --mode record --proxy-port 8081 ...
HTTP_PROXY=http://localhost:8081 pytest tests/integration/
## Step 3: Derive and validate contracts
curl -X POST "http://localhost:8080/_contracts/{group}?baseUrl=..."
## Step 4: Run mutation testing (11 strategies)
curl -X POST "http://localhost:8080/_contracts/mutations/{group}..."
## Step 5: Export and report
Report DONE if catch rate > 80%. BLOCKED if contracts fail.
Suggest: /ygs-fuzz-test for deeper coverage, /ygs-security-review for findings.
4. Speed at the Pipeline
The second knob from earlier simulation is pipeline speed, which is a cheaper lever. This section covers formal verification of the queue protocol itself, build/test optimization with real numbers.
4.1 Formal verification
Tests check specific inputs but formal verification proves properties hold over all possible inputs. For a concurrent system like a merge queue, the state space is too large to test exhaustively.
TLA+ specification of the merge queue
The spec lives at docs/examples/specs/merge_queue.tla and it models the core merge queue as a state machine:
Safety for Scope isolation
ScopeIsolation ==
\A s1, s2 \in Scopes :
s1 /= s2 => lanes[s1] \cap lanes[s2] = {}
Safety for Test before merge:
TestBeforeMerge ==
\A pr \in PRs :
prState[pr] = "merged" => testResults[pr] = "pass"
Liveness to ensure every PR eventually merges or is ejected:
Progress == \A pr \in PRs :
prState[pr] = "queued" ~> (prState[pr] = "merged" \/ prState[pr] = "ejected")
Dafny verification of merge invariants
The Dafny spec at docs/examples/specs/verified_merge.dfy proves four properties at compile time: batch merging preserves main branch health, bisection correctly partitions PRs, risk scoring is monotonic, and shard partitioning loses no tests:
predicate MainBranchHealthy(mergedPRs: set<PR>)
{
forall pr :: pr in mergedPRs ==> pr.testResult == Pass
}
lemma MergeBatchPreservesHealth(mainBranch: set<PR>, batch: set<PR>)
requires MainBranchHealthy(mainBranch)
requires forall pr :: pr in batch ==> pr.testResult == Pass
ensures MainBranchHealthy(mainBranch + batch)
{
// Proof is automatic: union of two sets where all elements
// satisfy the predicate still satisfies the predicate.
}
4.2 Test-impact analysis: scripts/mq/test_impact.py
The default behaviour is deliberate: always run the full test suite, partitioned into balanced shards. Diff-scoped runs are opt-in via --diff-scope and they require an explicit --head <branch> to define what to diff against.
# Default: full suite, 8 shards, no diff analysis
python -m scripts.mq.test_impact --pr-number main --num-shards 8
# Explicit diff-scope: numeric PR (base branch auto-extracted from GitHub API)
python -m scripts.mq.test_impact --pr-number 42 --num-shards 8 --diff-scope
# Explicit diff-scope: branch name (must supply --head as base to compare against)
python -m scripts.mq.test_impact --pr-number feature/my-branch \
--head main --num-shards 8 --diff-scope
The script performs following operation:
- Fetches changed files via
gh pr view --json - Detects language from file extensions (Python, Go, TypeScript, Java, Kotlin, Ruby, C#, Rust)
- Maps to test files using naming conventions
- Traverses imports one level deep to find transitive dependents
- Partitions into balanced shards using greedy bin-packing with historical timing data
- Falls back to the full suite if no tests map to changed files
4.3 Parallel fan-out: ai-parallel-test
Test-impact analysis feeds directly into Formicary’s fan-out for parallel shard execution (docs/examples/ai-parallel-test.yaml):
# analyze task: runs test_impact.py, emits ::add-job-context TestShards::
- task_type: analyze
script:
- python -m scripts.mq.clone_pr --pr-number {{.PRNumber}}
- python -m scripts.mq.test_impact --pr-number {{.PRNumber}}
# TestShards is now in job context — fan-out reads it directly
# Fan-out: one task per test shard, running in parallel
# 4 CPU cores + 16G memory per shard ? real intra-shard parallelism
- task_type: run-tests
container:
cpu_request: "4000m"
memory_limit: 16G
fan_out:
source: TestShards
item_var: shard
max_parallel: {{.MaxShards}} # unquoted integer — YAML strict types
fail_fast: false
script:
- python -m scripts.mq.clone_pr --pr-number {{.PRNumber}}
- python -m scripts.mq.run_scoped_ci --shard "{{.shard}}"
Each shard runs independently on its own Kubernetes pod. The polyglot test runner (scripts/mq/run_scoped_ci.py) detects the project type from marker files and builds the right command.
Here is a sample output:
## Test Impact Analysis
**191** / 191 tests selected (**0%** reduction) across **2** shards
## Test Results
? **87** / 87 passed
?? 130s wall clock across 2 shards
### Shard Performance
| Shard | Tests | Passed | Failed | Duration | Status |
|-------|-------|--------|--------|----------|--------|
| 1 | 44 | 44 | 0 | 130.5s | ? |
| 0 | 43 | 43 | 0 | 115.4s | ? |
> **Parallel speedup:** 246s sequential ? 130s parallel (1.9x across 2 shards)
### Slowest Tests
| Duration | Test |
|----------|-------------------------------|
| 30.01s | `TestInitEnabled` |
| 30.01s | `TestInitEnabled` |
| 10.02s | `Test_ShouldAsyncWithSleep` |
| 10.01s | `Test_ShouldAsyncWithSleep` |
| 4.68s | `Test_EncryptDecrypt` |
| 2.45s | `Test_ShouldRealGet` |
| 2.28s | `TestSharedSubscription` |
### Test Health Insights
- ? Pass rate: 100% (87 tests)
- ? Shard balance: 12% imbalance (115s – 130s)
- ? 4 slow tests likely using real sleeps — consider mocking time or
reducing timeouts: `TestInitEnabled` (30.0s), `Test_ShouldAsyncWithSleep` (10.0s)
- ?? 3 slow test names appear in multiple shards (possible test duplication):
`TestInitEnabled`, `Test_ShouldAsyncWithSleep`, `TestSharedSubscription`
4.4 Deployment risk
The merge queue optimizes everything up to the point where the commit lands on main but merged PRs still have to ship. The original simulation from Joe Magerramov’s formula models the probability that a merge batch is clean and continuous delivery. In my simulation I added support for batch PRs as it’s used in many organizations. But this also increases risk, e.g., a merge queue with a 95% batch success rate can still lead to a weekly release train that ships clean only 36% of the time. The formula extends naturally:
release_success = (1 - defect_rate) ^ prs_per_release
CD is the special case where prs_per_release = 1, e.g.,
| Defect rate | CD (1 PR) | Daily (10 PRs) | Weekly (50 PRs) | Bi-weekly (100 PRs) |
|---|---|---|---|---|
| 0.5% | 99.5% | 95.1% | 77.8% | 60.6% |
| 1% | 99.0% | 90.4% | 60.5% | 36.6% |
| 2% | 98.0% | 81.7% | 36.4% | 13.3% |
| 5% | 95.0% | 59.9% | 7.7% | 0.6% |
At 2% defect rate, a CD pipeline has a 98% clean deployment rate. The same team on a weekly train: 36%. Two-thirds of their releases ship with at least one defect.
The rollback trap
Suppose release R1 ships with a hidden bug. By the time monitoring detects it, R2 and R3 have already deployed. Rolling back to pre-R1 means reverting R2 and R3 too.
| Stacked releases | Rollback feasibility | Recovery strategy | MTTR multiplier |
|---|---|---|---|
| 1 | ? Straightforward | Rollback | 1.0x |
| 2-3 | ?? Costly but possible | Rollback (with intermediate revert) | 1.5x |
| ?4 | ? Impractical | Roll-forward (fix-forward only) | 2.5x |
One bad PR poisons the entire train and the release sits in an unknown state while someone bisects 50 changes to find the culprit.
Deployment maturity
Following are best practices from the industry for improving deployment maturity:
| Dimension | Weight | What it enables |
|---|---|---|
| Automated testing | 2.0x | CI on every PR catches defects pre-merge |
| Canary deployment | 2.0x | Progressive rollout catches prod-only failures |
| Automated rollback | 1.5x | Instant revert on anomaly detection |
| Observability | 1.5x | Alerting + metrics detect failures fast |
| Wave deployment | 1.0x | Stage ? preprod ? prod progression |
| Feature flags | 1.0x | Decouple deploy from release |
| Blue/green | 0.5x | Zero-downtime deploy infrastructure |
The maturity score translates into a risk multiplier — a low multiplier means your infrastructure catches most defects before they reach users:
adjusted_success = release_success + (1 - release_success) × (1 - risk_multiplier)
The MQ report now shows two critical markers:
- ? Blue line: your current position given your defect rate, batch size, and deployment maturity.
- ? Red line: the calamity threshold: the batch size where success drops below 70%.
The calamity threshold formula:
max_safe_batch = log(0.70) / log(1 - defect_rate)
The deployment risk model points to three levers:
- Reduce batch size. Move from weekly to daily trains.
- Invest in deployment maturity. Automated testing (weight 2.0x) and canary deployment (weight 2.0x) are the two highest-leverage dimensions.
- Prevent release stacking. Each additional stacked release multiplies recovery cost.
5. The Learning Loop
5.1 Multi-agent coordination
The merge queue is naturally an actor system where each PR is an actor with a lifecycle (queued ? batched ? testing ? merged/ejected). The properties that matter at scale:
- Location transparency: a PR doesn’t care which CI runner tests it
- Supervision: when a batch fails, the supervisor (bisect task) localizes the fault and recovers
- Isolation: scope-aware lanes are isolation boundaries
Better decomposition
Similar to the microservice architecture, the path forward is better decomposition:
- Smaller PRs
- Better module boundaries
- Formal verification
- Automated scope detection
- Dependent PR grouping, e.g., group related PRs so if one fails, the entire group is ejected.
5.2 The skills improvement flywheel

5.3 Metrics
The four DORA keys (deployment frequency, lead time, change failure rate, failed deployment recovery time) are table stakes. At agent-scale, you need additional metrics like:
- PR size
- PR pickup time Stale PRs conflict more often and carry higher bisection cost in the merge queue.
- AI PR acceptance rate.
- Rework rate.
- Cycle time decomposition break first-commit-to-deploy into pickup, review, merge, deploy sub-phases.
5.4 Path forward
The bottom line from all the evidence in this post:
- Don’t adopt AI at agent-scale before your delivery system can handle agent-scale quality.
- Pipeline speed is the cheaper lever. Defect rate has diminishing returns; pipeline duration has an order of magnitude of headroom.
- The merge queue is not a queue. It’s a coordination problem: scope-aware routing, risk-gated review, speculative batching, bisection-based fault recovery, and a learning loop.
- Formal verification catches what tests can’t. For concurrent protocols, tests check paths and TLA+ checks invariants.
- Specs reduce defects at source. Structured acceptance criteria, design docs for high-impact changes, and tiered review based on blast radius.
6. Getting Started
Everything above is code in three repos: formicary (workflow engine), ai-dev-tools (Python scripts), and you-got-skills (Claude skill library). Here’s how to test each piece.
Prerequisites
# 1. Deploy Formicary to Kubernetes (single-container: queen + ant + storage)
kubectl apply -f k8s.yaml
kubectl port-forward svc/formicary 7777:7777 19000:19000
# 2. Get an API token from the UI: http://localhost:7777 ? Profile ? API Token
export FORMICARY_TOKEN="<jwt>"
export GH_TOKEN="<github-pat>"
export SLACK_CHANNEL="your-channel"
# 3. Deploy all workflows + set configs
cd docs/examples
./deploy-ai-workflows.sh --create-k8s-secret --set-configs \
--gh-org bhatti --gh-repo todo-api-errors --bedrock
Test 1: Scope router
In Slack, mention the bot in your configured channel:
@bot scope main --repo https://github.com/org/repo
This prompt triggers ai-scope-router workflow and then creates a report showing the scope, blast radius, coverage, historical data for defects, security and other metrics.
Test 2: Parallel test with test-impact analysis
@bot parallel-test --repo https://github.com/org/repo main
@bot parallel-test --repo https://github.com/org/repo --branch develop
This prompt triggers ai-parallel-test workflow that fans out multiple shards for parallel tests and then creates a report showing the test results.
Test 3: Risk-gated review
@bot gate-review https://github.com/org/repo/pull/1
This prompt triggers ai-gate-review workflow and generates a report showing blast radius, security, test coverage and other metrics.
Test 4: Contract testing
@bot contract-test 1 --service myorg/my-api:latest
This prompt triggers ai-contract-test workflow and executes contract test against the target service. It then generates a report with the test results.
Test 5: Merge Queue Analysis
@bot mq https://github.com/org/repo
This prompt triggers ai-merge-queue workflow, evaluating health of the merge queue and calculating the risk/blast radius/complexity of all open PRs. Here is a sample snippet from the merge queue report:
? 250 open PRs across 22 scope lanes — repo: acme/platform
| Metric | Value |
|---|---|
| Total open PRs | 250 |
| Scope lanes | 22 |
| High-risk PRs | 140 |
| Needs human review | 148 |
| Conflict risk lanes | 1 |
?? 148 PR(s) flagged for human review before merge (high blast-radius, failed CI, or sensitive paths detected).
?? 1 lane(s) have conflict risk multiple PRs touching overlapping paths. Merge one at a time.
Overall status: ? Unstable (basis: 74.8% PRs aged >48h — CI status unavailable from API)
| Metric | Value |
|---|---|
| Total PRs in queue | 250 |
| CI failure rate | N/A — not reported by API (verify in CI dashboard) |
| PRs aged >48h | 187 (74.8%) |
| High blast-radius PRs | 140 |
| Medium blast-radius PRs | 82 |
| Approx batch size (PRs/lane) | 10.0 |
Unstable: 74.8% of PRs are aged >48h. Queue is filling faster than it drains — unblock stale PRs or shrink batch size.
| Category | PRs | Bug PRs | High-Blast | Hotspot |
|---|---|---|---|---|
| ? ui | 122 | 11 | ?? | ? yes |
| ? data | 32 | 3 | ?? | ? yes |
| ? api | 31 | 5 | ?? | ? yes |
| ? authn_authz | 31 | 5 | ?? | ? yes |
| ? security | 15 | 3 | ?? | ? yes |
| sre | 12 | 1 | ?? | — |
| backend | 3 | 0 | ?? | — |
| config | 1 | 0 | ?? | — |
| unknown | 3 | 0 | — | — |
Test 6: PR Audit
@bot pr-audit
This prompt triggers ai-gh-pr-audit workflow, evaluating the risk/blast radius/complexity of all merged PRs. Here is a sample snippet from the pr audit report:
| # | PR | Issue | Author | Title | Cat | Type | Blast | Risk | LOC | Files | Cx |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | #48401 | PROJ-44822 | jsmith | AuthCoordinator for REST OAuth | ? authn_authz | ? security | high | ? 72 | 1,847 | 23 | ? high |
| 2 | #47228 | AI-4931 | alee | fix(ai): allow regional Bedrock pr… | ? data | ? bug | high | ? 58 | 892 | 14 | ? medium |
| 3 | #48768 | DL-2875 | mchen | Implement more fine-grained IAM roles | ? authn_authz | ? feature | high | ? 45 | 634 | 11 | ? medium |
| 4 | #48909 | — | jdoe | SSRF guard for external MCP connector | ui | ? bug | medium | ? 32 | 312 | 8 | ? low |
| 5 | #48827 | PLAT-16031 | kpatel | [Flaky Test] fix race in bottlenec… | api | ?? refactor | low | ? 12 | 87 | 3 | ? low |
| … |
Related Reading
- What Happens After the PR Merges: Building the Learning Loop Software Factories Are Missing
- Orchestrating Background AI Agents for Software Teams
- Killing the State Machine: Declarative AI Coding Agents with an Orchestration System
- Building Production-Grade AI Agents with MCP and A2A
- Building a Production-Grade Enterprise AI Platform with vLLM
- Agentic AI for Personal Productivity: Building a Daily Minutes Assistant with RAG, MCP, and ReAct
- Agentic AI for Automated PII Detection with LangChain and Vertex AI
- Agentic AI for API Compatibility with LangChain and LangGraph
- AI Writes Code, You Own the Design
- Building a Distributed Orchestration and Graph Processing System (Formicary)
- RADAR: Automating Low-Risk Code Review at Meta
- From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI
- What I’m Hearing About Cognitive Debt (So Far)
- Software Factories in September 2026
- Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agents
- Agent Learning Flywheel: How AI Agents Improve
- A practical guide to risk-based code review
- 3,100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
- Minions: Stripe’s one-shot, end-to-end coding agents
- Running a Software Factory Efficiently at Uber Scale
- Spotify’s Background Coding Agent, Part
- Assess Risk with Blast Radius
- AI coding has made CI a bottleneck, so we reworked ours to keep up
- The Valley of Calm
- Merge Queues Were Built for Humans. AI Agents Need More.
- Merge Queues: how we ship faster with fewer incidents
- We Cut PR Merge Time by 92%
- The merge queue is the new bottleneck; DORA in the Age of AI
- 5 Ways to Stop AI Agents Stepping on Each Other
- AI coding made us faster. Why did incidents increase?
- Why change failure rates are rising 30%
- The 4-Minute Approval; Clean Code No Longer Signals Quality
- 2026 is becoming the year of AI quality
- AI Code Review Is Not Enough
- A Quality Gate for AI Coding Agents
- What is a Merge Queue
- Contract and Fuzz Testing Tutorial
- Accelerate State of DevOps Report
- Engineering Benchmark Report
- Octoverse 2024