Agentic delivery
Migrations in weeks, not yearsReview time down, not upEvery change verified, journaled, auditable

Ship Enterprise Software
at Agent Speed.
Keep the Control.

Coding agents made code cheap. Throughput, stability, and trust are still earned.
We build the harness — specs, verification, risk-routed review, governance —
that turns agents into a delivery system your engineers sign off on. Measured on your own metrics.

Your request becomes a clear spec

Techery has built and shipped software for

+ dozens of mid-market enterprises since 2007

Six Problems Where Agents Already Beat Headcount —
If the Harness Is Right

These are the shapes of work where agentic pipelines have produced citable enterprise results. They are also the work your best engineers dread, defer, and quietly budget a year for. We scope each one as a fixed outcome with a number attached.

Typical shape: Java 8/11 → 17, .NET Framework → .NET, AngularJS → React, Python 2 → 3, out-of-support SQL Server, RPG/COBOL comprehension.

Legacy modernization

Framework migrations nobody schedules because they touch everything and show nothing when they land.

80% of code in Google's largest migrations written by AI Google Research, 2024–25

Typical shape: Enzyme → RTL, JUnit 3 → 5, brittle end-to-end suites, characterization tests for code that has none.

Test-suite migration

Old test frameworks block every change, and thin coverage means nobody can prove a refactor is safe.

6 weeks for Airbnb to migrate 3,500 test files — not 1.5 years Airbnb Engineering, 2025

Typical shape: Fleet-wide library bumps, transitive-dependency CVEs, hardening findings from your last pentest.

Security & CVE fixes

Upgrade sweeps across hundreds of services and pentest findings that sit in the tracker for quarters.

4,500 developer-years saved by Amazon on Java upgrades Amazon, vendor-reported

Typical shape: Plain-English specs from legacy code, executable acceptance criteria, call-graph maps of a 15-year-old core system.

Tribal knowledge

Business logic lives in one engineer's head. We turn it into specs, ADRs and tests before they leave.

280K developer hours saved at Morgan Stanley in five months WSJ, 2025

Typical shape: PR size and review latency both rising, rubber-stamped merges, incidents traced to unread AI-generated code.

Review bottleneck

AI doubled your pull requests. Your senior reviewers didn't double — and they're burning out on diffs.

+91% more review time on teams with high AI adoption Faros AI, 2025

Typical shape: Integration glue, internal-tool backlogs, API adapters, recurring toil that never becomes a project.

Maintenance drag

The roadmap slips every quarter because the same team keeps the lights on. Agents take the toil.

72% of IT budget spent just keeping the lights on Forrester

AI Made Code Cheap.
It Didn't Make Delivery Faster.

In Every Engineering Org That Adopted Coding Agents Bottom-Up, the Same Four Things Broke — in the Same Order.

The teams that got the headline numbers — Google, Airbnb, Morgan Stanley, Nubank — didn't buy a tool. They engineered a system around it: context, evals, test harnesses, human gates, governance. That system is what we build. In your repos, on your metrics, owned by your team when we leave.

Output went up. Throughput didn't

PRs got 2.6× larger and waited 5.3× longer for a reviewer. Every 25% increase in AI adoption came with a 7.2% drop in delivery stability.

LinearB 2026 · Google DORA 2024

Review became the constraint

81% of developers now spend more time in code review than before AI. Senior engineers either rubber-stamp or burn out. One leaves, and the queue doubles.

Harness 2026 survey · The Pragmatic Engineer

Nobody could verify the output

Without a test harness, an agent's change is an opinion. 45% of AI-generated code carries an OWASP Top-10 flaw; Java fails 72% of the time.

Veracode GenAI Code Security Report 2025

The pilot never reached production

Over 40% of agentic AI projects will be cancelled by 2027 — unclear value, escalating cost, no risk controls. The tool worked. The system around it didn't exist.

Gartner, June 2025

How We Deliver

From Intent to Merged Change,
Every Step Verified

Five stages, three human gates, one journal. Humans decide what and approve the risky; agents do the typing; deterministic gates decide what counts as done. Questions and approvals reach people in Jira, Slack or the pull request — never in a terminal.

Human gate · spec sign-off

Intent

Ambiguity is removed before a single token is spent on code.

PRD grill session

The ticket and its linked docs are pulled from your tracker and wiki, then stress-tested with the owner until every open question has an answer or an explicit assumption.

Executable specification

  • BDD scenarios an agent must pass
  • Architecture Decision Records
  • Invariants that can never change

Risk class assigned

Auth, PII, payments, infra-as-code and data migrations are flagged as high blast radius from the start.

The loop: an open question goes back to the owner as a comment on the ticket. The spec is not signed off while a question is unanswered.

Human gate · plan approval

Plan

Research on the real code, then a plan a human approves.

Ground truth first

  • AST index and call graphs of the target modules
  • Existing tests, error paths, hidden contracts
  • Tracker history, linked docs, past PRs on the same files
  • Only the context the task needs — no repo dumps

Compressed implementation plan

  • Thin vertical slices, in order
  • Write scopes per slice — which files may change
  • Token and dollar ceilings per run

Fresh context for execution

Exploration history is discarded; agents start from the approved plan and the target files only.

The loop: a rejected or amended plan sends the agent back to research the gap and re-plan. Approval is a ticket transition — nothing executes before it.

Fully automated

Execute

Parallel agents, each in its own worktree, each returning a patch.

Orchestrator decomposes

Turns the plan into serialized sub-tasks with validation contracts. It never writes code itself.

Workers in isolated worktrees

  • Bounded concurrency, declared write scopes
  • Model routed per step — fast for boilerplate, deep for logic
  • Out-of-scope edits flagged or quarantined

Patches, not mutations

Nothing touches the integration branch until the workflow says so. A crash mid-run loses nothing.

The loop: every failed gate downstream comes back here — to the same worker, with the failure in context. Repair in-session, never a fresh guess.

Fully automated

Verify

Deterministic gates and a second model that tries to disprove the work.

Deterministic gates

  • Type-check, lint, architecture linters
  • Tests with hashed execution receipts — no faked green
  • Spec scenarios injected back on failure

Adversarial refutation

A separate agent on a different model tries to disprove each claim. Only findings that survive a majority vote proceed.

Security & blast-radius scan

OWASP checks, secrets, dependency policy, and the risk class re-evaluated on the actual diff.

The loop: a failed gate returns to the worker with the evidence. A claim that does not survive refutation is dropped and logged — never merged, never silently retried.

Human gate · high-risk only

Integrate

Risk-routed review, explicit merges, and a journal of every decision.

Risk-routed review

  • Low blast radius: holdout tests, then auto-merge
  • Medium: agent review + owner spot-check
  • High: human architect with the full reasoning trace

Explicit integration

Patches merge sequentially and in scope. Conflicts ask a person; they are never guessed.

Journal & report

  • Every prompt, tool call, patch and approval
  • Tokens and dollars per step
  • Checks, ledger, remaining risk

The loop: a merge conflict asks a person; failed post-merge checks reopen the build loop. Every decision — human or by policy — is in the journal.

Resume-safe end to end: a run survives a crash, a reboot, or an overnight wait on a human — and picks up from the exact step it stopped on.

Not Every Change Deserves
the Same Review

85%

of agent changes never wait
in a senior engineer's queue

Your architects read the 15% that can hurt you — with the full reasoning trace already attached.

Auto-merged or owner spot-checkArchitect sign-off

Every agent-authored change arrives with

Scoped patchPassing gatesRefutation verdictReasoning trace
Arrives

Blast-radius classifier

Deterministic rules first, model judgement second — mapped to your change-management policy.

  • Auth & identity
  • PII
  • Money movement
  • Infra-as-code
  • Data migrations
  • Public APIs
Routes

Low risk

No human touch

Holdout testsAgent reviewAuto-merge

Medium risk

Minutes, not days

Holdout testsAgent reviewOwner spot-check

High risk

Full reasoning trace

All gatesCross-vendor refutationArchitect sign-off

Fewer

Human reviews per merged change — the number we track first, because it is the one that was rising.

Smaller

Diffs — write scopes and vertical slices keep each patch reviewable in one sitting.

Stable

Change-failure rate held flat or better while throughput rises — the DORA definition of "actually faster".

Context From Your Systems In.
Questions to Your People Out.

The pipeline reads the ticket, the linked docs and the code history the way a good engineer would — and when it needs a human, it asks in Jira, Slack or the pull request, not in a terminal. Nobody on your side learns a new tool.

Context in

Issue & project trackers

JiraLinearAzure DevOpsGitHub IssuesAsana

Docs & wikis

NotionConfluenceGoogle DocsSharePoint

Code, CI & history

GitHubGitLabBitbucketYour CI

Chat & operations

SlackMicrosoft TeamsSentryDatadogPagerDuty
readsThe delivery pipeline

Exactly the context each step needs — nothing else

  • Least-privilege access
  • Sensitive fields masked
  • Every read logged
  • Anything with an API
asks

Your people, in their tools

Questions become tickets

Jira commentLinear sub-issueSlack thread
Developer focused at his screen while colleagues blur past

Developers stay
in the pull request

PR commentReview requestCODEOWNERS routing
Team reviewing printed reports around a table in warm evening light

Stakeholders get the report

Epic commentNotion status pageWeekly digest

Approvals
are transitions

Ticket statusSlack / Teams actionPR approval

How a question travels

Agents are not allowed to guess. An unanswered question suspends the run — durably, with nothing lost — and reaches the one person who can answer it, in their own tool.

01Agent

Hits ambiguity

Two plausible readings of the ticket. Guessing would be a bug.

02Pipeline

Suspends the run

State is journaled. Tokens stop burning. Other lanes continue.

03Pipeline

Posts the question

A comment or sub-task on the ticket, assigned to the owner. Slack pings them.

04Owner

Answers in-tool

In Jira, Linear or Slack — the way they already work. Hours later is fine.

05Pipeline

Validates the answer

Checked against the expected shape. A vague reply asks a sharper follow-up.

06Agent

Resumes

From the exact step it stopped on, with the answer — and its author — in the journal.

Works with

JiraLinearAzure DevOpsAsanaNotionConfluenceGoogle DocsSlackMicrosoft TeamsGitHubGitLabBitbucketSentryDatadogPagerDuty
…and anything with an API, including in-house trackers.

Brownfield Gets a Harness.
Greenfield Gets an Agent-Native Foundation.

Wrapping agents around a fifteen-year-old system and designing a new one so agents can build most of it are different jobs. Same engine, same principles, same gates — a different playbook, and a different first move.

Make the system you already
run agent-ready

No rewrite. Thin tests, undocumented logic and tribal knowledge become specs an agent can work from.

First outcome

A framework upgrade, test-suite port or CVE sweep — in one quarter, with numbers attached.

First move

Comprehend before touching anything

Architecture

Untouched first, agent-ready over time

Infrastructure

Your CI and tracker stay

Verification

Pin today’s behaviour first

Delivery

Strangler-fig slices, never big-bang

Autonomy earned in five steps
  1. 1Comprehend
  2. 2Pin behaviour
  3. 3Make agent-ready
  4. 4Slice & migrate
  5. 5Ramp autonomy

Both end in the same place:
an Enterprise Factory your platform team owns.

Brownfield gets there in slices. Greenfield starts there.

The Enterprise Factory:
Dark-Factory Throughput, Lights On Where It Matters

The problem

The "dark software factory" — agents producing code continuously with no human reading it — is real. Inside an enterprise it fails in predictable ways: architectural decay, faked test passes, silent divergence, no audit trail, a token bill nobody owns.

What we add

We keep the throughput and add what a regulated, multi-team organization needs: autonomy tiers by blast radius, deterministic gates, a journal an auditor can replay, and people placed exactly where the risk is.

Step 1

Work you already produce

Nothing new to file. The factory listens to what your organization already emits.

  • Tickets & epics
  • Executable specs
  • Incidents & alerts
  • CVE feeds
  • Flaky-test reports
  • Doc drift
Step 2

Classify by blast radius

Deterministic rules first, model judgement second. Every item is scoped before an agent sees it.

  • Risk tier
  • Write scope
  • Token & $ ceiling
  • Owner
Step 3

Three lines, by autonomy tier

Work lands on the line that matches its risk. A workload moves up a tier only on its own evidence.

Lights-out line

Runs continuously, no human in the loop

Dependency & CVE sweeps · flaky-test repair · docs & ADR sync · codemods · test backfill

0human touches
Supervised line

Humans approve plans, spot-check diffs

Feature slices · internal APIs · shared libraries · integration glue · bug fixes

1–2touches
Gated line

High blast radius, named sign-off

Authentication · payments · PII · IaC · schema migrations · public APIs

2–3named touches
Step 4

Same verification for every line

The tier decides who looks. It never decides whether the gates run.

  • Type-check & linters
  • Hashed test receipts
  • Refutation panel
  • Security scan
  • Budget check
Step 5

Merged, journaled, reported

Auto-merge on the lights-out line, explicit integrate elsewhere. The report posts to the epic; the journal keeps the rest.

  • Explicit integrate
  • Report to tracker
  • Immutable journal

What we keep from the dark factory

  • Continuous production, not sprints. Intent to merged change in hours or days on the lights-out line.
  • Agents do the typing. Humans write specs, approve plans and sign the risky 15%.
  • Specs and the journal are the durable assets. Code is cheap; what you own is what proves it.

What the enterprise version adds

  • Autonomy by blast radius. Three lines side by side; a workload moves up a tier only on its own evidence.
  • Deterministic gates and a replayable journal. Hashed test receipts, refutation, formal invariants where it counts; every decision recorded.
  • Governed identity, spend and infrastructure. Ephemeral agent credentials, budgets with breakers, continuous compute, a curated skills registry.

Diagnose. Build the Harness. Ship One Outcome. Hand It Over.

No transformation pitch, no big-bang cutover. Brownfield starts with a fixed-fee diagnostic that produces your baseline numbers; greenfield starts with a blueprint. Either way we ship one scoped outcome — and your platform team owns everything we build.

Same engine, gates and journal on both tracks. Only the starting point differs: an existing system or a new foundation.

3 weeksto Diagnose

Instrument before promising anything.

  • Baseline: review latency, PR size, change-failure rate, rework, lead time
  • Agent-readiness score per repo: types, tests, error clarity, docs
  • Knowledge-risk map: what only one person understands
  • Pick the wedge: the one outcome with the hardest ROI

You get: a numbers-first report and a scoped, priced first outcome. Useful even if you stop here.

8 weeksto Build the harness

The system the winners built in-house — in your VPC.

  • Executable specs, ADRs, architecture linters
  • Characterization tests where coverage is thin
  • Deterministic gates, hashed test receipts, budgets
  • Weft workflows wired to your git, CI and tracker

You get: a running pipeline your engineers can read, edit and test without us.

14 weeksto Ship the outcome

One migration, one test-suite port, one CVE sweep — with acceptance criteria written down first.

  • Strangler-fig slices, parity-validated, no big-bang
  • Risk-routed review live on real changes
  • Before/after on the diagnostic baseline

You get: the outcome, and a case study with your numbers in it — whether or not you let us publish it.

OngoingScale & hand over

Your platform team owns the harness. We stay only as long as it pays.

  • Lights-out lines for the low-risk work; gated lines for the rest
  • Second and third outcomes on the same pipeline
  • Embedded delivery engineers, month to month
  • Shared harness multiplier: fixes inherited by every team
  • Trace-mined evals so the pipeline keeps improving

You get: a delivery capability, not a vendor dependency.

2 weeksto Blueprint

Design the product so agents can build most of it.

  • Executable specs and domain boundaries as the source of truth
  • Architecture decisions and risk classes written down as ADRs
  • Delivery pipeline, gates and budgets designed with your team
  • Cost-per-feature model agreed before the first commit

You get: a blueprint your architects can challenge — and keep, whoever builds it.

5 weeksto Agent-native foundation

The infrastructure the hundredth feature will need, built before the first.

  • Repo scaffold: strict types, architecture linters, deterministic errors
  • Continuous compute: pre-merge sandboxes, hashed test receipts
  • Weft pipeline, journal, spend governors, skills and context registry
  • Tracker, wiki and chat wired in for context and human gates

You get: an empty product that already builds, tests and ships itself.

10 weeksto First slices

Vertical slices through the whole stack, at agent speed.

  • Agents do most of the typing; humans own specs and high-risk sign-off
  • Risk-routed review live from the first pull request
  • Cost, lead time and change-failure rate measured per feature

You get: a working product increment on the real pipeline, with its numbers.

OngoingScale & hand over

Your team runs the feature loop. We stay only as long as it pays.

  • Lights-out lines for the low-risk work; gated lines for the rest
  • Backlog to production on the pipeline your team already owns
  • Embedded delivery engineers, month to month
  • Trace-mined evals keep the harness improving as the product grows

You get: a product whose feature cost stays flat as it grows.

Eight Rules We Don't Negotiate

Every one of these exists because we watched a pipeline fail without it. They are the difference between a demo and a delivery system.

01

Spec before code

Natural language is where bugs are born. Intent becomes executable — BDD scenarios, ADRs, invariants — before an agent writes a line.

02

Deterministic gates over prompts

You cannot prompt a model into compliance. Type-checkers, linters and hashed test receipts reject what the prompt would only discourage.

03

Verification never trusts the author

The agent that wrote the change never grades it. A separate validator, on a different vendor, tries to refute every claim.

04

Humans are steps, not spectators

Approvals live inside the graph and suspend the run durably. An answer can arrive hours later and is validated like any other output.

05

Patches, not mutations

Agents work in isolated worktrees with declared write scopes. Nothing lands until the workflow says integrate.

06

Every run is journaled

Prompts, tool calls, patches, approvals, spend — appended to an immutable log. Any run can be replayed, audited, or resumed.

07

Measure throughput, not output

Lines of code and PR counts are vanity. We report review latency, change-failure rate, lead time and human touches per outcome.

08

You own the harness

Model-agnostic, runs in your VPC, no proprietary runtime. When we leave, your platform team keeps everything — including the source.

The Gains Are Real. So Are the Failure Modes.

We cite both, because the counter-evidence is the strongest argument for doing this with a harness instead of a seat license. Sources are independent or clearly labelled.

With a harness around the model

80%of code changes in Google's largest migrations authored by the pipeline; a two-year effort halved.Google Research · peer-reviewed experience report
6 wksto migrate ~3,500 test files at Airbnb, against a 1.5-year manual estimate. Retry loops, not heroics.Airbnb Engineering blog, 2025
280k hdeveloper hours saved at Morgan Stanley in five months, translating 9M lines of legacy code into readable specs.Wall Street Journal, June 2025
42%faster task completion in ANZ Bank's controlled trial of 100 engineers — one of the few enterprise RCTs.arXiv / The Register, 2024

Without one

−19%experienced developers were slower with AI tools on mature repos — while believing they were 20% faster.METR randomized controlled trial, July 2025
−7.2%delivery stability for every 25% increase in AI adoption. AI amplifies the org you already have.Google DORA 2024 · reaffirmed in DORA 2025
45%of AI-generated code introduced an OWASP Top-10 vulnerability across 100+ models; Java failed 72% of the time.Veracode GenAI Code Security Report, 2025
40%+of agentic AI projects will be cancelled by end-2027: escalating costs, unclear value, inadequate risk controls.Gartner, June 2025

Cited industry results, not Techery client claims. Every engagement produces its own numbers; we publish them only with client sign-off.

We Run Every Engagement on Weft.
Then We Leave It With You.

Weft is our open-source engine for durable, journaled, schema-validated multi-agent coding workflows. It isn't the product — the delivery system we build in your repos is. But it's why that system survives crashes, audits, and vendor changes.

Open source · MIT · github.com/techery/weft
.weft/workflows/audit-and-fix/main.tsabridged
// find bugs → refute across vendors → human gate → fix in worktrees → integrate → test
export default defineWorkflow({
  description: "Find bugs, verify across vendors, fix with approval",
  input:  z.object({ paths: z.array(z.string()) }),
  output: z.object({ fixed: z.array(Finding), skipped: z.array(Finding) }),
}, async (ctx, { paths }) => {
  ctx.phase("Find");
  const found = ctx.successes(await ctx.parallel(paths, (p) =>
    ctx.agent(`Find correctness bugs in ${p}. Cite file:line.`, {
      schema: z.object({ findings: z.array(Finding) }),
      key: `find:${p}`,
    })));
  const findings = found.flatMap((r) => r.findings);     // typed — no nulls

  ctx.phase("Verify");                                    // a different vendor grades
  const real = ctx.successes(await ctx.pipeline(findings)
    .step((f) => ctx.parallel(["claude", "codex", "claude"], (provider, i) =>
      ctx.agent(`Try to refute: ${f.claim} (${f.file}:${f.line})`, {
        schema: Verdict, provider,
        key: `refute:${f.file}:${f.line}:${i}`,
      })))
    .filter((votes) => votes.filter((v) => v.real).length >= 2)  // majority
    .map((_votes, f) => f)
    .run());

  ctx.phase("Fix");
  const go = await ctx.gate({                               // a human, durably
    action: `Apply fixes for ${real.length} bugs`, risk: "medium",
  });
  if (!go.approved) return { fixed: [], skipped: real };

  const fixes = ctx.successes(await ctx.parallel(real, (f) =>
    ctx.agent.detailed(`Fix and add a focused test: ${f.claim}`, {
      schema: FixResult,
      isolation: "worktree",                       // own tree, returns a patch
      write: { paths: [f.file, "**/*.test.ts"], mode: "warn" },
      key: `fix:${f.file}`,
    })));
  const ledger = await ctx.integrate(fixes, {
    order: "sequential", onConflict: "ask",
  });

  ctx.phase("Check");
  await ctx.check("tests", { exec: ["pnpm", "test"], required: true });
  return {
    fixed: real.filter((f) => ledger.merged.includes(`fix:${f.file}`)),
    skipped: [],
  };
});

What makes a workflow durable

  • The graph is your code

    Workflows are ordinary TypeScript. await is a sequential edge, parallel fans out, if on a typed field is a branch. No DSL to keep in sync.

  • Every step returns a schema-validated value

    Agent turns, human answers, shell commands, git reads — all typed. Invalid output is repaired in-session with the errors fed back, never silently accepted.

  • Journaled, resumable, edit-tolerant

    Append-only event log. A run survives a crash or a reboot and resumes from the step it stopped on. Rewording a prompt re-runs only what depended on it.

  • Humans are steps

    ctx.gate, ctx.human.ask, approve, review suspend the run durably. The answer can come hours later — from the CLI, the web UI, a Claude Code / Codex session over MCP, or the Jira comment and Slack thread we wire in for your team.

  • Patches, scopes, explicit integrate

    Write steps run in their own git worktree and return a patch. Out-of-scope files are flagged or quarantined; nothing lands until ctx.integrate().

  • Two vendors, one interface, real accounting

    Claude and Codex behind one provider API with per-step routing. Tokens and dollars enforce hard ceilings shared with sub-workflows.

  • Testable without a model

    Fixtures match on step keys and go through the same schema validation — a workflow with human gates and shell checks runs end to end in a unit test.

Weft plugs into your stack

Your environment
GitHub / GitLabYour CIJira / LinearNotion / ConfluenceSlack / TeamsYour VPC

Weft — the delivery engine

Workflows as code. Journal, gates, worktrees, budgets, checks, report.

  • CLI
  • Web UI
  • MCP server
  • Daemon
Claude Agent SDKCodex SDKMock for tests
Model providers

Your code never trains a model.
Zero-retention provider agreements by default; air-gapped deployment with self-hosted models where policy requires it.

Three Companies Call Us. They Have Different Problems.

Two of them are living with a system that exists; one is about to build a new one. All three are too small for a global integrator's minimum engagement and too under-resourced to turn $20 agent seats into a system. If one of these reads like your Monday, we should talk.

You bought speed. You got a review queue.

The Velocity Trap

A product or tech-enabled company, 60–250 engineers. AI assistants arrived bottom-up 12–18 months ago. PR volume doubled. The DORA metrics didn't move — except change-failure rate, which went the wrong way.

  • Senior reviewers are the bottleneck, and one has already left
  • A post-mortem has named AI-generated code
  • The board has asked what the AI spend delivered
  • Your platform team is 2–6 people and underwater

Start with: A 2–4 week instrumented diagnostic of review latency, PR size and stability — then the harness and risk-routed review, owned by your platform lead.

The system works. The person who understands it is retiring.

The Legacy Rebuild

A mid-market company — distribution, manufacturing, insurance, specialty finance — with a 12–25-year-old core on .NET Framework, IBM i/RPG, Delphi or a dead vendor's fork. Undocumented logic. An IT team staffed for operations, not engineering.

  • A dated trigger: end of support, an audit finding, a retirement notice
  • A previous rewrite failed or was shelved after the quote
  • A business sponsor — CFO, COO or operating partner — not just IT
  • Willingness to fund comprehension before committing to a program

Start with: A fixed-fee comprehension phase that extracts the business logic into specs and tests, quantifies key-person risk — then modernization in parity-validated slices. No big-bang cutover.

You're building something new. Build it so agents can build it.

The New Build

A funded product team, a company launching a new platform, or a rewrite that has already been decided. A blank repo, a deadline, a small team — and the choice of whether the first hundred features ship the way the last hundred did.

  • A greenfield repo or a decided rewrite, with a date on it
  • A team of 3–30 that wants leverage, not headcount
  • No entrenched CI or review culture to fight
  • A CTO who wants cost per feature to stay flat as the product grows

Start with: A two-week blueprint — executable specs, domain boundaries, ADRs, risk classes — then the agent-native foundation and the first vertical slice on it.

Questions Your CTO and CISO Will Ask

Code custody, models, measurement,
and what we need from you — answered plainly.

Does our source code leave our environment?

The pipeline runs in your VPC or on your runners; the journal, patches and worktrees live in your git. Model calls go to the providers you already approve — Claude or OpenAI under enterprise, zero-data-retention terms — and only the context a step needs is sent, never the repository. Where policy forbids any external call, we deploy air-gapped with self-hosted models. Nothing you own trains a model.

Which models and tools do you use? Do we have to switch?

Weft is model-agnostic: Claude and Codex sit behind one provider interface, routed per step, and a cross-vendor panel grades findings so no single model marks its own homework. Your existing Copilot, Cursor or Claude Code seats stay — the pipeline runs alongside them, and engineers can drive it from inside those sessions over MCP.

How do you measure whether it worked?

Baseline first, promises second. The diagnostic instruments PR review latency, PR size, change-failure rate, rework and lead time before we change anything. Each outcome is then scoped with acceptance criteria and a number: hours, cost, coverage preserved, defect escape rate, human touches per merged change.

You keep the dashboards. The engagement report is generated from the journal, not written after the fact.

Our test coverage is poor. Can you still do this?

Yes — and we start there. Below roughly 60–70% coverage on the target modules, agent output is unverifiable, so the first outcome becomes a test harness: characterization tests that pin current behaviour, generated and reviewed at scale. That harness is what makes the migration safe, and it is the single most reused asset we leave behind.

What do you need from our team?

A named owner on your platform or DevEx side — they own the harness when we leave, so they are in it from week one. Repository, CI and tracker access. Two to four hours a week from the owners of the target modules. And early access to whoever holds the undocumented knowledge; their cooperation is a hard dependency, and we design the engagement so they become its champion, not its casualty.

Do our product managers and stakeholders need to learn a new tool?

No. The pipeline plugs into the tracker, wiki and chat you already run — Jira, Linear, Azure DevOps, Notion, Confluence, Slack, Teams — and works in both directions. It reads the ticket, its comments and linked docs for context. When an agent needs a decision, the question appears as a comment or sub-task assigned to the right person; approvals are a status change or a button; progress and the final report post to the epic; developers get review requests on the pull request.

The Weft web UI and CLI exist for your platform team, who own the pipeline. Everyone else keeps working where they already work.

We're building something new from scratch. Does this apply, or is it only for legacy?

It applies — with a different playbook. Brownfield work fits a harness around a system that exists; greenfield work designs the system so agents can build most of it. That means an agent-native architecture (modular boundaries enforced by linters, strict types, service contracts as schemas, error messages a model can act on) and infrastructure to match: continuous-compute sandboxes instead of classic CI, hashed test receipts, the journaled pipeline, spend governors and a skills registry — all in place before the first feature.

Greenfield engagements start with a one-to-two-week blueprint instead of a diagnostic, and the first vertical slice typically ships within the first quarter on the pipeline your team keeps.

Is this the "dark factory" idea — code nobody reads?

It borrows the throughput and rejects the blindness. A dark factory sets one autonomy level for everything and stops reading the code; the failure modes are well documented — architectural decay, agents faking green tests, divergence nobody notices for months. The Enterprise Factory runs three production lines side by side: a lights-out line for work a test can prove (dependency sweeps, flaky-test repair, docs, codemods), a supervised line for features where a person approves the plan, and a gated line for anything with real blast radius where a named architect signs with the full trace.

Every line goes through the same deterministic gates and the same journal. A workload moves up a tier only when its own numbers say it can — never by mandate.

Are you replacing our engineers?

No. Typing was never the bottleneck. Your engineers move to the work that decides outcomes: writing specs, approving plans, reviewing the high-risk 15% of changes with a full reasoning trace in front of them. Their review load goes down because low-risk changes stop reaching them. Every skill the pipeline needs — harness engineering, context curation, spec writing — is taught during the engagement.

How does this hold up under SOC 2, SOX, or a regulator?

Every step is journaled to an append-only log: prompts, tool payloads, patches, test receipts, approvals and spend, with secrets redacted. Auto-approvals by policy are recorded as decisions, not omitted. Risk classes map to your change-management policy, so high-blast-radius changes always carry a named human sign-off. Any run can be replayed for an auditor. We provide a controls one-pager mapped to SOC 2 and OWASP to shorten your security review.

How long until we see something real?

The diagnostic takes two to four weeks and is useful on its own. The first shipped outcome typically lands within the first quarter. Modernization programs are staged in parity-validated slices — old and new run side by side, and nothing is a big-bang cutover.

What if the pilot doesn't hit its number?

Outcomes are fixed-scope with acceptance criteria written down before work starts, so "hit" is not a matter of opinion. If it misses, you still keep the harness, the specs, the tests and the source — they are yours either way, and they are the parts that were expensive to build.

Why not Accenture, EPAM, or the agent vendor's own deployed engineers?

Global integrators sell multi-year programs with minimums that don't fit a $200M–$3B company. Vendor deployed-engineering teams are excellent at making their product stick — and leave the harness, verification and review design to you. We are 200+ senior engineers who have shipped production systems for enterprises since 2007 and production AI since 2021, sized for a first outcome in a quarter, and indifferent to which model you run.

Book a Delivery Diagnostic

Two to four weeks, fixed fee. You get a baseline on review latency, change-failure rate and agent readiness — and one scoped outcome with a number attached. Useful even if you stop there.