Back To All

The artifact chain is the measurement chain

October 1st, 2026
Topics
Agents
AI
AI ADOPTION
AI Agents
AI IMPACT
AI SDLC
Share Article

The artifact chain is the measurement chain

The ADLC framing going around LinkedIn this week gets the structure right: every stage commits an artifact, the next stage reads it, and the chain of commits becomes the audit trail. That is a bigger idea than it looks, and it is not an audit trail until somebody actually writes the files.

Six stages, one loopPlan, design, build, test, deploy, maintain, closing back on itself
Artifact in, artifact outEvery stage commits something the next stage reads
Humans above the loopReviewing what the agent flagged, rather than writing every line

Alex Circei, CEO of Waydev  /  September 2026  /  10 minute read

Rakesh Gohel posted a breakdown this week that has done about 750 reactions and 110 reposts, arguing that the SDLC is dead and the ADLC, the Agentic Development Lifecycle, replaces it. Anthropic uses the term AI-native SDLC for roughly the same shift.

The argument is clean. Every gate in the traditional lifecycle existed because writing code was slow and expensive, so we built rituals to force alignment across the weeks or months of build work. When agents write most of the diff, the build phase collapses to hours and the premise behind the rituals disappears. His line for the consequence: “The difference between SDLC and ADLC isn’t speed. It’s structure.”

The stage-by-stage version, briefly. Plan becomes an intent.md written by the originator in their own words, version-controlled and machine-actionable from the first commit. Design and requirements collapse into one session with an agent, guided by policy encoded as skills. Build turns institutional knowledge into CLAUDE.md files the agent reads every session, with guardrails running as hooks rather than habits. Test replaces stage-gate QA with continuous evals, where each session verifies its own work before a human sees it. Deploy layers agentic review and reserves human judgment for critical code. Maintain closes the loop: a production alert writes a new intent and flows back through the pipeline with no human in the invocation path.

I want to pick up the part that I think is being underrated, which is not the loop. It is the artifacts.

Every artifact is an instrument

For the fifteen years I have been around engineering measurement, the central problem has been that intent was never written down anywhere machine-readable. We had commits, and we had tickets that described work in whatever vocabulary a team happened to use, and everything else was inference. You could measure that a change happened. You could rarely measure whether it was the change somebody meant to make.

If the ADLC is implemented properly, that changes. Intent becomes a file. Design becomes a file. The plan becomes a file. Review becomes a record rather than a conversation. The production signal points back at the intent that produced it. For the first time the lifecycle is natively instrumented, not because anybody set out to instrument it, but because agents need artifacts to work from and humans need them to supervise.

The artifact, and the thing it lets you measure intent what to build Intent to production lead time spec design Rework traced to spec defects plan approach Scope drift between plan and diff code + tests the diff Agent-authored share and autonomy level review record approval Flag precision and review latency production signal incident, metrics Escape rate and loop closure time The loop closes here Under the old lifecycle, only two of these six existed as machine-readable artifacts. Everything upstream of the diff lived in tickets, documents and people’s heads, which is why every engineering metric of the last decade started at the commit and worked backwards by inference.

Figure 1. The artifact chain, read as an instrumentation chain. The measurements in the lower row are mine rather than part of the original framing, but each one becomes available the moment the artifact above it exists.

Intent used to be the thing we inferred. In this model it is the first commit.

The chain is only as strong as its weakest file

Here is my problem with how this is being received. The diagrams circulating describe an end state, and people are reading them as a description of where agentic teams already are. In almost every organization I see, the chain has two links and a gap where the other four should be.

How often the artifact actually exists, in the organizations I see intent.md rare, and usually a ticket renamed spec exists for big work, absent for the rest plan lives in a chat session that nobody kept code + tests always review record an approval click, rarely a record of reasoning production signal exists, but not linked back

Figure 2. Illustrative, from what I see in the field rather than from a survey. The two links that were always there are still the two links that are there. Agents did not create the upstream artifacts by arriving.

A partial chain is worse than no chain, because it looks like an audit trail without being one. If the intent file is a ticket title copied into a markdown file to satisfy a template, then every downstream measurement inherits that emptiness. You get a beautifully structured record of nothing in particular, and the first time somebody asks why a change was made, the answer is still going to be a Slack search.

The discipline the model demands is not tooling. It is that somebody writes down what they actually wanted, in enough detail that a machine can act on it and a human can later be held to it. That is the same discipline good specs always required, and most organizations were never good at it. Agents raise the cost of being vague, because vagueness now compiles.

Who evaluates the flagger?

The other claim worth examining is that humans do not leave the loop, they move above it, going from writing every line to reviewing what the agent flagged.

I agree with the direction and I think the sentence hides the entire problem. Reviewing what the agent flagged is only a safe position if the flagging is good. The quality of that flagging is now the load-bearing property of the whole system, and I have not yet met an organization that measures it.

What the agent flags Flags everything Humans review all of it anyway, then stop reading carefully. Alert fatigue Flags what matters Human attention lands on the changes that actually carry risk. The only quadrant that works Flags noise, misses risk Worst of both. Reviewers are busy and unprotected. Flags rarely, misses things Feels efficient. Defects reach production and nobody connects them back. Low precision Low recall Flags too much Flags too little Two numbers tell you which quadrant you are in: what share of flags turned out to matter, and what share of escaped defects were never flagged.

Figure 3. Agentic review quality, treated as a measurable property rather than an assumption. Moving humans above the loop transfers the burden of judgment onto the flagger, so the flagger needs an accuracy record.

Two metrics cover most of it. Flag precision: of the things the agent raised, what fraction did a human agree needed action. Flag recall, which is harder and more important: of the defects that reached production, what fraction were never flagged at all. Precision protects attention. Recall protects the system. Almost everyone will measure the first and skip the second, because the second requires connecting incidents back to the review that let them through, and that link is exactly the one the artifact chain is supposed to provide.

Loop closure time is the metric this model deserves

The maintain stage is the boldest part of the framing: a production alert writes a new intent and flows back through the pipeline with no human in the invocation path. Set aside for a second whether you want that in your environment. It creates a measurable interval that did not exist before, and it is a better number than anything we currently use.

Call it loop closure time. The clock starts at the production signal and stops when a fix has been verified in production, with the intent, plan, diff and review all linked in between. It is mean time to recovery with the causal chain attached, which means that for the first time you can ask not just how long recovery took but which stage consumed the time.

ArtifactWhat it unlocksWhat breaks without it
intentLead time measured from the moment somebody wanted the thing, not from the first commitYou measure delivery from the point work became visible, which flatters every team with a long queue in front of it
specRework attributable to specification defects, separated from rework caused by implementationAll rework looks like engineering sloppiness, and the upstream cause never gets fixed
planScope drift, measured as the distance between the approach agreed and the diff producedAgents expand scope silently and nobody notices until review, which is the most expensive place to notice
code and testsAgent-authored share by service, and autonomy level per changeAdoption is a licence count and autonomy is a guess
review recordFlag precision, flag recall, review latency, and who actually approved whatGovernance is asserted rather than evidenced, and moving humans above the loop is an act of faith
production signalEscape rate tied back to authorship, and loop closure time end to endIncidents and changes live in separate systems, so the feedback loop never actually closes

What I would do about it

  1. Start with intent, not with tooling. The cheapest high-value change available to most teams right now is requiring a short written statement of what is wanted and why, committed before the agent session starts. No template gymnastics. If it cannot be written in a paragraph, that is information too.
  2. Make the links machine-followable. Every artifact needs an identifier that the next one carries. This is unglamorous plumbing and it is the difference between a chain and a pile. Without it, none of the delivery measurements in that table are computable.
  3. Measure your agentic reviewer before you trust it. Sample flagged and unflagged changes, have humans grade both, and track precision and recall over time. A reviewer you have not evaluated is not a control, it is a filter of unknown shape.
  4. Instrument loop closure time now, even partially. Signal to verified fix, with whatever links you have. The gaps in the measurement will tell you exactly which artifact to invest in next.
  5. Treat the agent-facing files as production assets. Skills, hooks and the context files agents read every session shape every downstream artifact. They deserve ownership, review and change history at least as strict as the code they generate.

Where I would temper the framing

Declaring the SDLC dead makes for a good headline and slightly misdescribes what is happening. The six stages are not disappearing, they are changing cost, ownership and duration. Plan, design, build, test, deploy and maintain all still occur in the agentic version, which is why the same six words appear in both diagrams. What changed is that the expensive stage became cheap, the cheap stages became expensive, and the handoffs became artifacts instead of meetings. That is a genuinely large change. It is a restructuring, not a funeral, and teams that treat it as a funeral tend to throw out the gates they still needed.

The part of this I find most encouraging has nothing to do with speed. For a decade, everyone measuring engineering has worked from the commit backwards, reconstructing intent from evidence that was never meant to carry it. A lifecycle where intent is written down first, in a file, under version control, is a better world for anyone who cares whether the work matched the plan.

But that only holds if the files are real. An artifact chain full of placeholder files is an audit trail that will pass an audit and teach you nothing. The structure is the opportunity. The discipline to fill it is still the job.

Source: Rakesh Gohel, post on LinkedIn, September 2026, introducing the ADLC framing and the six-stage breakdown summarised here, and referencing Anthropic’s use of the term AI-native SDLC. The stage descriptions and the artifact chain concept are his. Figures 1 to 3 and the measurement mappings throughout are mine; the proportions in Figure 2 are illustrative of what I see in practice rather than survey data.

Ready to unlock your SDLC productivity?

Request a Demo Call