Back To All

Harness Engineering: The AI SDLC Is Won Outside the Model

September 10th, 2026
Topics
AI SDLC
Harness AI SDLC
Share Article

Harness Engineering: The AI SDLC Is Won Outside the Model

Every engineering leader I talk to is having the same argument, and it is the wrong one.

The argument is about models. Which one writes the best code. Which one benchmarks highest on the latest eval. Whether to standardize on one vendor or let teams choose. Meanwhile the teams that are actually getting value out of AI in their software delivery lifecycle have quietly stopped caring. They swapped their model three times last quarter and barely noticed, because the model was never the hard part.

The hard part is the harness.

AAAWgmp1bWIAAAAeanVtZGMycGEAEQAQgAAAqgA4m3EDYzJwYQAAABZcanVtYgAAAEdqdW1kYzJtYQARABCAAACqADibcQN1cm46YzJwYTo3YjFkYmIyNi0zMjdhLTRmNTEtYmY5ZS0yMGY5M2VmZDZmMDkAAAADl2p1bWIAAAApanVtZGMyYXMAEQAQgAAAqgA4m3EDYzJwYS5hc3NlcnRpb25zAAAAALxqdW1iAAAARGp1bWRjYm9yABEAEIAAAKoAOJtxE2MycGEuaW5ncmVkaWVudC52MwAAAAAYYzJzaNYm2VrvNgkglc+agqVGM1IAAABwY2JvcqNpZGM6Zm9ybWF0bWltYWdlL3N2Zyt4bWxqaW5zdGFuY2VJRHgseG1wOmlpZDozZWU3MzU0Yy1lMjYzLTQzNjktODNiMS1hY2U5MWI3MDE2ZTVscmVsYXRpb25zaGlwaHBhcmVudE9mAAAB4mp1bWIAAABBanVtZGNib3IAEQAQgAAAqgA4m3ETYzJwYS5hY3Rpb25zLnYyAAAAABhjMnNoShlbXHQZ9SrTUZ8DkoYrSQAAAZljYm9yomdhY3Rpb25zgqJmYWN0aW9ua2MycGEub3BlbmVkanBhcmFtZXRlcnOha2luZ3JlZGllbnRzgaJjdXJseC1zZWxmI2p1bWJmPWMycGEuYXNzZXJ0aW9ucy9jMnBhLmluZ3JlZGllbnQudjNkaGFzaFgg2cwvaXssnytWOKeukVyq94+WOK7g0TzkZQuRTHjQgpukZmFjdGlvbngdY29tLmFudGhyb3BpYy5jbGF1ZGUucHJvdmlkZWRqcGFyYW1ldGVyc6F4H2NvbS5hbnRocm9waWMub3JpZ2luLWNvbmZpZGVuY2VndW5rbm93bmtkZXNjcmlwdGlvbnhmQ2xhdWRlIHByb3ZpZGVkIHRoaXMgZmlsZSBhdCB0aGUgcmVxdWVzdCBvZiBhIHVzZXIgYW5kIG1heSBoYXZlIGNyZWF0ZWQgb3IgbW9kaWZpZWQgdGhlIGZpbGUgY29udGVudHMubXNvZnR3YXJlQWdlbnShZG5hbWVmQ2xhdWRlcmFsbEFjdGlvbnNJbmNsdWRlZPUAAADIanVtYgAAAEBqdW1kY2JvcgARABCAAACqADibcRNjMnBhLmhhc2guZGF0YQAAAAAYYzJzaHTcMaSpsV625TopT5TyV4EAAACAY2JvcqVjYWxnZnNoYTI1NmNwYWRNAAAAAAAAAAAAAAAAAGRoYXNoWCC695KbtyCZzyDo3ecBidAPs2Ubl3xbHqt2v06a4FC44mRuYW1lbmp1bWJmIG1hbmlmZXN0amV4Y2x1c2lvbnOBomVzdGFydBjYZmxlbmd0aBkeBAAAAj5qdW1iAAAAJ2p1bWRjMmNsABEAEIAAAKoAOJtxA2MycGEuY2xhaW0udjIAAAACD2Nib3KlY2FsZ2ZzaGEyNTZpc2lnbmF0dXJleE1zZWxmI2p1bWJmPS9jMnBhL3VybjpjMnBhOjdiMWRiYjI2LTMyN2EtNGY1MS1iZjllLTIwZjkzZWZkNmYwOS9jMnBhLnNpZ25hdHVyZWppbnN0YW5jZUlEeCx4bXA6aWlkOjMwYzFlZDBhLWE1NDEtNDY5My1iMjJjLWU4MGU5ZGU2ZThhNnJjcmVhdGVkX2Fzc2VydGlvbnODomN1cmx4LXNlbGYjanVtYmY9YzJwYS5hc3NlcnRpb25zL2MycGEuaW5ncmVkaWVudC52M2RoYXNoWCDZzC9peyyfK1Y4p66RXKr3j5Y4ruDRPORlC5FMeNCCm6JjdXJseCpzZWxmI2p1bWJmPWMycGEuYXNzZXJ0aW9ucy9jMnBhLmFjdGlvbnMudjJkaGFzaFgg5SXdrx0YahoEPHrM1m/DQkklXsXuLHk3xpv/oMqr7sSiY3VybHgpc2VsZiNqdW1iZj1jMnBhLmFzc2VydGlvbnMvYzJwYS5oYXNoLmRhdGFkaGFzaFggM5KOjAsoD4v+2M5zc305ICtt5gefKq5316DVkQr28Zh0Y2xhaW1fZ2VuZXJhdG9yX2luZm+jZG5hbWVvQW50aHJvcGljIEZpbGVzZ3ZlcnNpb25lMS4wLjBrc3BlY1ZlcnNpb25lMi40LjAAABA4anVtYgAAAChqdW1kYzJjcwARABCAAACqADibcQNjMnBhLnNpZ25hdHVyZQAAABAIY2JvctKEWQISogEmGCFZAgowggIGMIIBjaADAgECAhRA5aAK7sI50L64g/oGQgU9Z1UTADAKBggqhkjOPQQDAzBJMRcwFQYDVQQKEw5BbnRocm9waWMsIFBCQzEuMCwGA1UEAxMlQW50aHJvcGljIENvbnRlbnQgQ3JlZGVudGlhbHMgUm9vdCBDQTAeFw0yNjA4MDcxODQzNTZaFw0yODA4MDYxOTQzNTZaMEQxFzAVBgNVBAoTDkFudGhyb3BpYywgUEJDMSkwJwYDVQQDEyBBbnRocm9waWMgQ2xhdWRlIENvbnRlbnQgU2lnbmluZzBZMBMGByqGSM49AgEGCCqGSM49AwEHA0IABJh6CmvLUBgFFNU0vUKlOVtE6djd17L5SuwX0LemFisBM3dkd/3cyjxFA3Qo5S46fX0/ihY0VZ7mfb9KF703t5OjWDBWMA4GA1UdDwEB/wQEAwIHgDAVBgNVHSUEDjAMBgorBgEEAYPoXgIBMAwGA1UdEwEB/wQCMAAwHwYDVR0jBBgwFoAUzlHiBIFOZFsj+OPEz5o+nMHXXMIwCgYIKoZIzj0EAwMDZwAwZAIwMXMdFJ4BetLLVY7ORuE9noqbbAZOZn/aArXyTwFAZfKrPzxF2vPoJNf1+UCdg1XGAjBwX1zd9WGqYkqmL5SFqw1QySjr1zJfpJM9+1rdDwSPLMOPOjKuiXjoU/pUUeG9RwmhY3BhZFkNngAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAPZYQGSBsJnB25zR/jvjDdK3E/5cAkmo8O5slDPBki45GU9BHfLLvrPHaCHeAGY6nrPjl5YKg6hpu4B/Hl5XQ0L2Z48= The Harness Drives the Value in the AI SDLC Models write code. The harness turns that into software you can ship, trust and measure. Coding Models Interchangeable engines. Powerful, and replaceable. General model Broad reasoning over the codebase Fast model Inline completion, low latency Reasoning model Refactors and root-cause analysis Code-specialised model Diff generation and repair Review model Second opinion on risky changes Swap models as needed. The harness works with any of them. The Delivery Harness Orchestration. Control. Integration with how you actually ship. Context Memory Repo access Permissions Workflow Guardrails Evaluation Routing Observability Integration Agent Turns model output into changes that reach production. Plan Break the work into steps Write Generate the diff Verify Tests, checks, review Ship Merge, deploy, watch Engineering Outcomes What the business actually buys the AI for. Reliability Changes that hold up in production Codebase knowledge Grounded in your repos and history Automation Toil handled end to end Developer experience Less waiting, less rework Trust Reviewable, compliant, auditable Speed Shorter cycle time to value Measured impact Adoption, impact and ROI you can report Dependable AI, measurable delivery impact. model capability shipped work Connected to Your SDLC The harness wires the agent into the systems your team already ships with. Git repos CI/CD Issues IDEs and agents Code review Engineers Same harness. Any model. Real repos. Real releases.
Models provide the intelligence. The harness turns it into software that ships, holds up, and can be measured.

What a harness actually is

A model is an engine. It produces intelligence on demand: a diff, a plan, an explanation, a test. What it does not produce is a change that lands safely in your codebase, passes your checks, respects your permissions, and can be explained to an auditor six months later.

The harness is everything between the engine and the outcome. Context assembly. Memory. Repository and data access. Permissions. Workflow. Guardrails. Evaluation. Routing between models. Observability. Integration with the systems your team already uses. The agent sits inside that structure and does the work: plan, write, verify, ship.

Put differently, the model provides capability. The harness converts capability into shipped software. Two organizations can license the exact same frontier model and get results that differ by an order of magnitude, and the entire delta lives in the harness.

That framing has a practical consequence. If the harness is the converter, then value is not produced in one step. It is produced along a chain, and a chain has links that break.

The conversion chain

The Conversion Chain: Six Places Capability Turns Into Value Every link needs a harness layer. A missing layer does not slow the chain down, it breaks it. 01 Intent to task Workflow, task shaping Turn a ticket or a request into work an agent can actually take on. Without it The agent solves the wrong problem, precisely and at speed. 02 Task to grounded prompt Context assembly, memory Pull the code, decisions and conventions that bear on this change. Without it Plausible code that ignores how your system is actually built. 03 Prompt to candidate change Routing, model access Send the work to an engine sized for it, at a cost that makes sense. Without it Shallow reasoning on hard work, or a premium bill for a rename. 04 Candidate to verified change Guardrails and evaluation Prove the change is correct and safe before a human ever looks at it. Without it Review queues absorb the whole gain. Seniors become graders. 05 Verified to shipped Permissions, CI/CD Move it down the same path every other change in your org takes. Without it Work stalls at the edge of what the agent is allowed to reach. 06 Shipped to measured value Observability, telemetry Attribute outcomes back to the work, in terms the business recognises. Without it No evidence of impact, and no second year of budget. Capability enters on the left. Value only exists on the right. A model upgrade improves the input to link 03. It does nothing for links 01, 04, 05 or 06.
Capability enters on the left. Value only exists on the right. Each link depends on a specific harness layer.

Six things have to happen for model capability to become business value, and only one of them is affected by which model you licensed.

Intent to task. Somebody wants something. That has to become work an agent can take on: scoped, unambiguous, with a definition of done. Skip this and the agent will solve the wrong problem precisely and at speed, which is worse than not attempting it, because now somebody has to review the wrong thing.

Task to grounded prompt. The model has to know your architecture, your conventions, the reason the previous team wrapped that service in a retry loop, and which of the four similar utility functions is the sanctioned one. None of that is in the weights. Skip it and you get plausible code that ignores how your system is actually built.

Prompt to candidate change. Now the model runs. This is the only link a vendor upgrade improves.

Candidate to verified change. Generation is cheap and getting cheaper. Confidence that the generated thing is correct is expensive and does not scale automatically. Skip verification and review queues absorb the entire gain, with senior engineers turned into graders of plausible code written by nobody.

Verified to shipped. The change has to travel the same path every other change in your organization takes, with the same permissions, approvals, and deployment controls. Skip this and work piles up at the edge of what the agent is allowed to reach.

Shipped to measured value. Somebody eventually has to answer what changed, why, on whose authority, and whether it worked. Skip this and there is no evidence of impact, which in most companies means no second year of budget.

Five of the six links are harness engineering. This is why teams with mediocre model access and excellent harnesses beat teams with the reverse, consistently, and why the gap widens rather than narrows as models improve.

Context assembly is the most underrated layer

If I had to point at one layer that separates the demos from the deployments, it is context.

Context Assembly: The Most Underrated Layer in the Harness The context window is a budget, not a bucket. What you spend it on decides the quality of the output. Where the knowledge lives Repositories Current code and structure Pull request history Why things changed Design docs and ADRs Decisions and constraints Tickets and specs What is being asked for Incident history What already went wrong Runtime and telemetry How it behaves in production The assembly pipeline 1 Recall Search broadly across sources for anything that could bear on the task. 2 Rank and filter Score by relevance, recency, ownership and blast radius. Most of it gets cut. 3 Fit the budget Decide what survives the window. Summarise the rest rather than dropping it silently. 4 Assemble Compose one task packet: goal, constraints, code, prior decisions, expected tests. The task packet What the agent actually sees The goal, stated once The code that will change The code that must not Conventions and prior decisions Tests and acceptance criteria Permission scope for this task What to do when unsure Agent Same model. Same prompt style. Radically different output, because the input is grounded. Retrieval is the easy half. Deciding what to leave out is the engineering.
The context window is a budget, not a bucket. Deciding what to leave out is the engineering.

The knowledge an agent needs is scattered across systems that were never designed to be read together: repositories hold the current state, pull request history holds the reasoning, design docs hold the constraints, tickets hold the intent, incidents hold the scar tissue, and telemetry holds the truth about behaviour in production.

Assembly is a pipeline, not a search box. Recall broadly, then rank by relevance, recency, ownership and blast radius, then fit what survives into a budget, then compose a task packet. The ranking step is where most of the quality is created and where most implementations are weakest, because pulling more context is easy and pulling the right context is not.

Two things are worth saying plainly here.

The first is that the context window is a budget, not a bucket. Bigger windows raise the ceiling on what you can include, but the cost of including irrelevant material is not zero. Noise dilutes attention, raises spend, and increases the chance the agent anchors on the wrong precedent. A disciplined harness summarises what it cannot afford to include rather than dropping it silently.

The second is that context quality is an organizational asset, not a technical one. Teams with clear ownership, readable pull requests, current documentation and honest postmortems get better AI output from the same model than teams without. The harness can only assemble what exists. This is the least fashionable reason AI programs underperform and one of the most common.

Autonomy is a grid, not a switch

The second layer where organizations get stuck is permissions. The usual approach is a single global decision: agents suggest, or agents open pull requests, or agents merge. One setting, applied to everyone.

Autonomy Is Not a Setting, It Is a Grid The harness decides how much freedom an agent gets, per class of change, not per user. Default on Gated by checks Human decides Human decides Default on Default on Gated by checks Human decides Default on Default on Default on Gated by checks Default on Default on Default on Default on Docs, tests, tooling Isolated service Shared library, schema Core or regulated systems Blast radius of the change Merge behind a flag Open a PR when all checks pass Draft a PR for a human Suggest in the editor Autonomy granted to the agent Most organisations pick one row for everything. That is either too slow to be useful or too loose to be safe. The grid is the point: the same agent can be trusted in one column and supervised in another.
Autonomy should vary by blast radius. The same agent can be trusted in one column and supervised in another.

That single setting is always wrong in one of two directions. Set it conservatively and the agent is useless for the routine work where it would have paid for itself. Set it aggressively and you have handed write access to your most sensitive systems to a process nobody is watching closely.

The better model is two dimensional. Autonomy on one axis, blast radius of the change on the other. Docs, tests and internal tooling can run at high autonomy from day one, because the cost of being wrong is low and the feedback is immediate. Shared libraries and schemas sit in the middle, gated by checks. Core and regulated systems keep a human in the decision, not as a formality but because that is where the asymmetry of outcomes is severe.

This grid is also how you expand safely. You do not raise autonomy globally. You promote one cell at a time, when the evidence from the cell below supports it. That gives you an expansion policy rather than an argument.

Routing and the economics nobody models

Not every task deserves a reasoning model, and not every task survives a fast one. Renaming a symbol across a service and untangling a distributed race condition are different problems with different cost, latency and failure profiles.

A harness that routes work to an appropriate engine spends less and fails less. The routing signal is usually available before the work starts: change size, file sensitivity, test coverage in the affected area, whether the task is mechanical or exploratory, whether a previous attempt failed. Teams that route on those signals typically find that a small fraction of tasks needs their most expensive engine, and that the same fraction is where most of the value sits.

The corollary matters for procurement. If your harness routes properly, model pricing becomes an optimization rather than a strategic commitment, and vendor lock-in stops being a board-level concern. Anything model specific belongs behind an interface. If switching providers requires touching agent logic, prompts, evaluation, and integrations all at once, the harness is doing its job badly.

Verification is the new bottleneck

Here is the pattern I see most often, and it is worth naming because it looks like success right up until it does not.

Adoption metrics are excellent. Pull request volume climbs. Then review latency stretches, rework creeps up, and change failure rate moves in the wrong direction. Nothing broke. The bottleneck simply moved, which is what always happens when you accelerate one stage of a pipeline and leave the rest alone.

The response is not to slow generation down. It is to make verification scale at the same rate: automated tests that actually gate, static and security analysis in the loop rather than after it, evaluation suites that catch regressions in agent behaviour, and progressive delivery so that being wrong is cheap and reversible. Human review then gets spent where judgment is genuinely required, instead of on the volume that machines should have caught.

An agent that verifies its own work before a human sees it is worth several agents that do not.

Three failure modes worth recognising early

The pilot that never scales. A few strong engineers get remarkable output because they are personally acting as the harness. They know what context to paste, when to distrust the output, and which changes to throw away. The value is real and completely non-transferable. Roll it out to two hundred engineers and it evaporates.

Volume without verification. Covered above. The tell is that adoption metrics look better every month while delivery metrics do not move.

The model treadmill. Every few months the organization re-evaluates vendors, re-runs bake-offs, and re-litigates the standard. It feels like diligence. It is displacement activity, because the model was never the constraint, and the work that would actually move the number keeps getting deferred.

Measure the outcome, not the activity

If the harness is where value is created, then the harness is where measurement belongs. Most programs measure the two easiest things and then struggle to defend their budget.

The Measurement Ladder: Most Programs Stop Two Rungs Too Early Each rung is real data. Only the top one answers the question your CFO is asking. Adoption Who has access and who turns up Seats assigned, weekly active users, tools enabled per team Tells you nothing about whether the work got better. rung 1 of 4 Activity What the AI produced Suggestions accepted, AI-authored diffs, agent sessions, tokens spent Volume is not value. This rung is where vanity metrics live. rung 2 of 4 Flow What happened to the delivery system Cycle time, review latency, rework rate, change failure rate, WIP Now you can see whether the gain survived the pipeline. rung 3 of 4 Business outcome What the organisation got Capacity redeployed, time to value, cost per change, incident load, ROI The only rung that renews the budget. rung 4 of 4 Harder to measure, harder to argue with If you can only report rungs one and two, you are describing usage, not impact.
Each rung is real data. Only the top one answers the question the CFO is asking.

Adoption tells you people have access. Activity tells you the tool produced output. Neither tells you the organization is better off. Flow metrics are the first rung where the delivery system itself is in view: cycle time, review latency, rework, change failure rate. Business outcomes are the rung that renews the budget: capacity redeployed, time to value, cost per change, incident load, return expressed in terms finance already recognises.

The questions worth answering are narrower and harder than the ones most dashboards answer:

  • Where is AI actually being used, by which teams, on which kinds of work
  • Does AI assisted work move faster from first commit to production, or just from idea to pull request
  • What happens to review load, rework, and change failure rate as AI volume grows
  • Which harness layer is the current constraint on all of the above
  • What is the return, expressed in terms your CFO already recognizes

This is the part most organizations skip, and it is the reason so many AI programs stall at the pilot stage. Not because the technology underdelivered, but because nobody could show what it delivered.

Where to start

If you are early, the sequence that works is boring and it is the same one every time.

Pick one workflow with real volume and low blast radius. Instrument the current baseline before you change anything, because a baseline you reconstruct afterwards is an argument, not evidence. Build the context assembly for that one workflow properly rather than generically. Put verification in front of human review rather than after it. Give the workflow one owner. Then read the flow metrics, not the usage metrics, and promote one cell of the autonomy grid when the evidence supports it.

That is a quarter of work, not a week. It is also the only version of this I have seen survive contact with a large engineering organization.

The uncomfortable implication

Harness engineering is not glamorous. It is context plumbing, permission models, evaluation suites, routing logic, and integration with the tools your team already lives in. It does not trend. It will never be a keynote moment.

It is also the entire ballgame. Models are converging and commoditizing, and their capability is available to your competitors on the same terms it is available to you. What is not available to them is your codebase, your context, your controls, and the system you built to turn a general purpose engine into dependable output inside your specific organization.

That system is the differentiator. Everything else is an engine you rent by the token.


Alex Circei is CEO and co-founder of Waydev, an AI-native engineering intelligence platform that measures AI adoption, impact, and ROI across engineering organizations.

Ready to unlock your SDLC productivity?

Request a Demo Call