AI SDLC: Who Owns "Done"?

Agents write the code and the tests. Who defines what "done" means? A quality engineering view of the AI SDLC: the bill, the roles, and the loop.

AI SDLC diagram from Anthropic's AI-Native SDLC Playbook. Before agents, all six stages (plan, design, build, test, deploy, maintain) run at human speed and build takes the longest. After agents, build runs at agent speed and shrinks to a sliver, while plan, design, test, deploy, and maintain keep their human-speed length, leaving cycle time to reclaim.

Last week I was in a leadership meeting when the VP of Engineering repeated the question his CFO had asked him the day before:

"If code is no longer a bottleneck, why do we need a team this size?"

The question is legitimate, and top tech firms and platforms are answering it with more or less the same playbook: give agents more autonomy, keep more of them running, and get people out of the loop.

At the same time, a second instruction is landing on those same teams: "We need to use AI to ship more, but without breaking production or leaking something we can't take back." Funny, isn't it? Ship faster, better and safer, all at once.

An AI SDLC is a software development lifecycle where agents do most of the implementation and people shift their attention to the gates: intent, specs, verification, and release. The open question for most teams is who defines the criteria those gates check against.

If you lead an engineering team moving to an AI SDLC, this post will help you navigate its challenges and get the most out of it. To get there, let's first take a look at what the companies building the next generation of tools, harnesses and infrastructure for this "AI era" are saying.

Building an AI SDLC? Abstracta Intelligence brings quality context, guardrails, and evidence brings quality context, guardrails, and evidence into the AI SDLC.
Curious about how to create “Release Readiness" Score?

The Industry's Answer: Move Human Attention to the Gates

In August 2026, Anthropic published its AI-native SDLC playbook (others call it the agentic SDLC, or Agentic Development Life Cycle, ADLC), describing the shape most teams are landing on: "human attention shifts along with the artifacts". People stop reviewing line by line and start reviewing at gates.

Days later, Kiro (Amazon's agentic IDE) published the most explicit version of that playbook I've read so far, including the results it reports from Amazon's own teams. Their opening line is the thesis: software development has split between those who changed how they work with agents and those who only changed their coding tools.

Kiro's framework has ten principles, but these four are the ones I found most valuable:

  1. Maximize agent time, minimize your involvement. Stop prompting and waiting. Give the agent a long task with validation built in, let it loop for thirty minutes or overnight, run several in parallel, review asynchronously. The constraint on your output stops being how fast you type and becomes how many agents you can keep meaningfully busy.
  2. Build your codebase for agents and give them a fast feedback loop. This is shift-left taken seriously. Linters, unit tests, local mocks, the full stack on a laptop, property-based tests derived from the spec, and the agent has access to all of it so it can find and fix its own failures before a person ever looks at the code. Their warning is one I'd underline: faster generation without this investment only means more broken builds.
  3. Invest in agent context. Steering files, architectural decisions written down, inline commentary kept as persistent memory. The teams that skip this step, they say, keep wondering why the agent repeats the same mistakes.
  4. Trust the boundaries, not the agent. Constrain what the agent can touch (files, tools, network, credentials), gate only what you can't undo, layer deterministic checks such as SAST and secret scanning, and then let it run unsupervised inside those limits.

Both documents are right about a lot, and the benefits are real. Getting started even looks easy: I strongly recommend you this set of scaffolding skills for each step, that you can also customize for your codebase, then a harness to orchestrate them without losing control and to give each specialized agent the right context.

What neither document defines is who verifies the loop, and by what criteria. They tell you where human attention goes; they don't tell you who exercises it or what "good enough" means for your system. That gap is the one quality engineering has to close, and it's what the rest of this post is about.

Comparison of what the AI SDLC playbooks from Anthropic and Kiro define and what they leave open. Defined, where human attention goes: review at gates instead of line by line, constrain what agents can touch, gate only what you can't undo, layer deterministic checks such as SAST, secrets and tests, and give agents context and a fast feedback loop. Left open, who exercises that attention and how: who verifies the loop, by what criteria a change is good enough, what counts as evidence, who owns the service that decides what gets checked, and where the learnings go. Quality engineering closes that gap.

🤝 Where I'd Sign

About seventy percent of it. "Trust the boundaries, not the agent" is our positioning in different words and "make intent explicit before code" is what we've been calling Intent Review. Kiro itself says this is an engineering investment and not a switch you flip, and that the first weeks feel slower. That's why we insist on a foundation phase before scaling agents. And "measure correctness, not just speed" is a sentence I could have written.

The remaining 30% sits underneath those ten principles, and it's where I think the real leverage is. Let’s drill down on them with these glasses: the bill, the roles, and the loop.

🔍 What's Underneath

1. The Bill: This Is Capex, and it Scales with What You Already Operate

Code became cheap. The bill didn't disappear; it moved, and it has three line items.

The first is the foundation. Every one of the ten principles asks you to build something before you get anything back: steering files, a codebase agents can navigate, property tests, boundaries, evals. Most teams I met skipped this and bought the licenses instead.

The second grows with your output. Anthropic just published its own numbers, and they tell the story better than I can:

Anthropic's own numbers on agentic coding in the AI SDLC: engineers ship 8x as much code per quarter as in 2021–2025, Claude authors 80% of that code, the codebase has 10x more tests, and CI jobs grew 25x in six months, which overloaded the test selection service. Three patches followed, a bigger machine that held 70 days, sharding that held 29 days, and daily restarts that held less than a day, before a redesign that took one engineer three weeks. Source: Anthropic, "Agentic coding is straining CI", September 2026.

The third is time. Kiro says the first weeks feel slower; behind a security review and a change board, weeks become quarters. And agent time only compounds if the gates answer at the same speed: an agent that finishes at 2am and waits until Monday isn't maximizing anything.

How big the bill gets depends on what you already operate. A startup born with agents has nothing in production to protect. Anyone running a legacy core, a bank being the extreme, changes a system whose rules live in the heads of the people who keep it working. Same playbook; different sequence, cost and risk.

So I'd answer the CFO in his own language. Treat the foundation and the verification infrastructure as capex, amortized over every release that follows. Treat the agents' work on features as opex, measured per change you actually accept. Who staffs each line?

2. The Roles: Developers Own the Harness, Testers Own "Done"

The frontier pages describe the developer's shift well, and I'd push it one rung further:

Writing code → reviewing code → reviewing the AI's review → improving the harness that produces the code.

The industry has a name for that top rung now: loop engineering. Design the cycle the agent runs (act, verify against real signals, decide, repeat, stop) instead of writing every instruction by hand. Kiro's tenth principle, "continuously tune your agent setup", is that job, and it's permanent.

Now look for the tester in the same pages. Testing appears as "the tests the agent writes". Domain experts and quality engineers are not in the picture. Anthropic's CI story has a telling detail: the service that decides what gets verified had, in their words, murky ownership. The thing that decides what gets checked had no owner.

That's a problem, because the agent that writes the code also writes the tests, runs them and declares success. The producer is the auditor. OpenAI just showed where that ends: during internal evaluations, an agent exploited a flaw in its own testing interface to copy the reference solution and collect the reward, and agents reasoned about the grader instead of the task.

One of the top labs in the world, and its response was graders that check how a task was done, not only whether. In our work with agents we treat it as a rule: the agent's self-assessment is never evidence. Only deterministic checks or a human validation count.

Every bank already knows this principle as segregation of duties, and Anthropic's playbook applies it at the pull request, where the agent that wrote the code can't approve it. What it leaves open is who decides what the checks should verify in the first place.

This is where I think the frontier framing falls short: very often the tester holds more business knowledge than anyone else on the team. Functional testers spend their days understanding the business logic and what users actually need; when they aren't there, teams end up running UAT with the users, which tells you where the knowledge lives.

In banking, testing is still the bottleneck, and the usual fix is "give me another tester". That scales execution. Agents scale execution better. What doesn't scale is the judgment: which module touches refunds, which rule is regulated, which failure would matter. Anthropic's test selector decides from history and package relevance, which is the right start; someone still has to say what matters if it fails.

So if there has ever been an era where shift-left is possible, it's this one. Everyone talks about designing the loop. We ask who designs the verification inside it: the criteria, what counts as evidence and when to escalate. That is the tester's ladder, and the top of it is a quality engineer:

Executing tests → designing tests → designing the verification inside the loop → owning the criteria and the evidence that make a release defensible.

Two ladders in the AI SDLC. Developer ladder: writing code, reviewing code, reviewing the AI's review, improving the harness that produces the code. Tester ladder: executing tests, designing tests, designing the verification inside the loop, owning the criteria and evidence that make a release defensible.

Developers own the harness; quality engineers own what "done" means."

Those last two rungs are the answer to the CFO. The team doesn't get smaller because agents write code. It changes shape, because someone has to own what "done" means.

3. The Loop: Shift-Left Now Includes Learning

Shift-left used to mean testing earlier. With agents it reaches the intent and the spec, before a line of code exists: a quality engineer reviewing an intent for ambiguity, a risk analysis attached to a spec, acceptance criteria and properties written before the agent starts. What moves left is the learning, not just the testing. Every criterion made explicit at the intent is one less thing to discover in a diff.

Anthropic's other piece of advice in the CI post: instrument your services so they become the agent's eyes and ears. Read it from the quality seat and it's the biggest upgrade testers have had in years. The same signals that let an agent hill-climb on a service let a quality engineer see what changed, what it touched and how it behaved last time in production, while testing, not after.

We saw this in a core banking system with hundreds of thousands of programs: turning system traces into plain language gave testers a view of the execution path behind each operation that nobody had while testing before, and verification got more precise because of it. Observability stops being the SRE's dashboard and becomes the evidence layer the criteria are graded against.

The AI SDLC loop from a quality engineering view: intent, spec, code, verify, release. Learning moves left when criteria, risk analysis and properties are written at the intent and spec, before the agent starts. After release, learnings come back into the harness, a skill, or the agent's context when the next decision is made, with an author and a date attached. Today most learnings evaporate at the end of a sprint.

Then there's the half nobody designs for. Everything above produces learnings: a criterion a tester made explicit, a boundary that turned out too wide, a module that failed for a reason nobody had written down. Today most of them evaporate at the end of a sprint. They become doubly valuable the moment you can put them back into the harness, a skill, or the context an agent receives, at the moment the agent needs them. A learning is only worth something once it lives where the next decision is made.

That's the direction we're taking with Abstracta Intelligence: quality knowledge that doesn't stay in one head or one sprint, but returns to the loop when it's needed, with an author and a date attached. It deserves a post of its own, and it's my next one.

🧾 The Platforms Are Seeing the Same Thing

Almost at the same time as Anthropic's playbook, the CEO of GitLab, published "When code is abundant". His argument comes from the platform side: once implementation becomes cheap, the unit of economics changes. It stops being code produced and becomes the cost of a change you can actually accept. And the line I keep coming back to:

"When implementation becomes abundant, trust becomes scarce." — Bill Staples, CEO of GitLab

GitLab isn't stating this as a principle; it's the direction of their product. When the company that runs your pipelines decides that acceptance cost, not commits, is the number to optimize, the argument for verification stops being a quality team's opinion.

I'd only add the team's side of the same coin. Cost per accepted change measures what you pay. What you buy with it is trust. So next to every agent-time dashboard I'd put a much simpler question: of everything your agents produced this week, how much did you accept without rework, and could you show the evidence behind it to your regulator, your CISO and your customers?

At Abstracta, we sum it up in a sentence: quality is how we engineer trust.

Illustration of a person at a laptop placing a chess piece on the screen and holding a connected-nodes icon, representing strategy and thoughtful decision-making. Faqs section about AI SDLC.

FAQs about the AI SDLC

What Is an AI SDLC?

ADLC is not the traditional SDLC with AI in the loop, it is a different process redesigned with SOTA LLMs performance and agentic AI in mind. In that model, the primary execution unit is an agent, and the vast majority of the implementation work is done by agents, and people focus on the gates that matter most, such as intent, specs, verification, and release. The familiar phases of planning, design, build, test, deploy, and maintain still exist, but they now run as a continuous feedback loop.

How Is the AI SDLC Different from the Traditional Software Development Life Cycle?

The traditional software development life cycle is human-driven. Work moves to the next stage through handoffs between different teams, and people review most artifacts line by line. In an AI SDLC, agents write the code and the tests, each stage leaves a versioned artifact in version control (such as intent and spec .md files), and human review concentrates at the gates. The build phase shrinks, work flows in multiple directions, and planning, verification, and release become the new bottlenecks. Most organizations need to update their mental model of the development process to match today's reality and put the investment there.

What Happens to Engineering Teams when AI Coding Agents Write Most of the Code?

When AI coding agents take over code generation, routine tasks, and most repetitive tasks, software engineering teams take on a different shape. Developers move from writing to reviewing code, then evolving to review AI’s review, and then toward improving the harness and the loop that produces it. Testers move from executing and designing tests toward designing the verification inside that loop and owning the criteria and evidence that make a release defensible. For Fabián Baptista, cofounder and CIO @Abstracta, the team pays for itself in the risk it keeps off the release.

Is Human Review Still Needed in an AI-Driven SDLC?

Human review is still needed in an AI-driven SDLC, and it works best when it concentrates where judgment matters. Line-by-line code reviews stop scaling once agents produce most of the diff, so human oversight moves to specific gates: approving intent and specs, changes to production systems, anything that touches sensitive data, and decisions that only humans should make because they can't be undone. Security teams, product owners, and quality engineers keep the final call at those gates, deterministic checks handle the routine verification in between, and human intervention is reserved for what autonomous systems shouldn't decide alone. At Abstracta, we design those gates with each team, starting with what can't be undone in their context. This is responsible AI applied to software delivery.

What Role Does Testing Play in the AI SDLC?

The role we've been arguing for years testing should have played: quality engineering, the discipline that defines what "done" means and how it gets proven. The AI SDLC doesn't create that need; it removes the excuse. Agents can generate and run tests faster than any team, so executing tests stops being where testers add value and owning the criteria becomes the job. The opportunity is bigger than it looks, because agents also let testers do things they couldn't before: read a code change and understand what it touches, follow observability signals and traces in plain language, and see how a module behaved in production while testing, not after. That is shift-left reaching the intent and the spec, and it lands on the people who already hold the most business knowledge on the team. At Abstracta, we work with testing teams to make that move, from executing tests to owning the criteria and evidence behind each release.

How Does Spec-Driven Development Fit into the AI SDLC?

Spec-driven development is the point where shift-left reaches the intent. Before an agent writes any code, the team turns product ideas, user feedback and stakeholder input into a spec the agent can read: business intent, requirements, edge cases and acceptance criteria. AI can draft that spec from what's already written down. What isn't written down is the problem, and in legacy systems that's most of it: the rules live in the heads of the people who keep the system working. That's why at Abstracta we add an Intent Review, where a quality engineer checks the intent for ambiguity, asks the people who hold the rules, and attaches a risk analysis before the agent starts. Every criterion made explicit at the intent is one less thing to discover in a diff.

What Should Engineering Teams Build before Scaling AI Agents?

Before scaling AI agents, engineering teams need a foundation in every part of the workflow where agents operate: context and steering files (CLAUDE.md in Claude Code, steering files in Kiro), architectural decisions written down, a codebase agents can build and test locally, property-based tests derived from the spec, clear limits on what agents can touch, and evals that check the agents themselves. Kiro's warning is the one to underline: faster generation without this investment only means more broken builds. The first weeks feel slower, which is why we treat the foundation as capex, amortized over every release that follows, rather than a switch you flip when you buy the licenses.

How Should Teams Measure the Impact of AI on Software Delivery?

,Teams should measure the impact of AI on software delivery by the changes they actually accept, not by the volume agents produce. Useful metrics: cost per accepted change, the share of agent output accepted without rework, total verification cost per release before and after agents, and whether the evidence behind each release would hold up in front of a regulator, a CISO or a customer. Agent-time dashboards and speed metrics are still useful, but only next to those. Velocity measured without quality tells you how fast you're producing things you may not be able to ship.

How Can Teams Keep What They Learn from Getting Lost between Sprints?

Teams can keep what they learn from getting lost by putting each learning back where the next decision is made. A criterion a tester made explicit, a boundary that turned out too wide, or a module that failed for a reason nobody had written down should return to the harness, a skill, or the agent's context, with an author and a date attached. That is the direction of Abstracta Intelligence, our enterprise AI platform built on Tero, our open-source agentic framework, so quality knowledge stays available for continuous improvement.

How Can Abstracta Help Teams Adopt an AI SDLC?

Abstracta helps engineering teams adopt an AI SDLC by owning the part most playbooks leave open: the criteria, verification and evidence that decide whether agent output is ready to ship. We start with a foundation phase, then work alongside developers and quality engineers to design the verification inside the loop, supported by Abstracta Intelligence and Tero, our open-source agentic framework. With nearly 20 years in quality engineering and experience in complex environments such as core banking, we help teams move faster while keeping control of what reaches production. It's the same shape as the answer to the CFO: with a top bank in Latin America we tripled QA coverage, from 6 to 18 projects, without linear headcount growth. The team didn't get smaller. It changed shape.

What I'd Take from All This

The bill. Maximize agent time, yes, and treat the foundation and the verification as the capex they are. The agents' work is opex, measured per change you actually accept.

The roles. Developers own the harness and the loop. The tester's ladder ends in a quality engineer who designs the verification inside that loop and owns the evidence, because the producer can't be the auditor.

The loop. Shift-left now includes learning: criteria at the intent, observability as the evidence layer, and what you learn coming back to where the next decision is made.

That's how "ship faster, better and safer" stops being three competing instructions. And it's the answer to the CFO that holds: the team pays for itself in the risk it keeps off the release, not in the code it no longer types.

If your team is already running agents overnight: who wrote the criteria they're graded against, who checked the tests they wrote, and where do the learnings go in the morning to make ADLC better each day?

At Abstracta, we help engineering teams answer those three questions with evidence. We work alongside developers and quality engineers to define the criteria agents are graded against, design the verification inside the loop, and bring what each release teaches back into the agents' context, supported by Abstracta Intelligence and Tero, our open-source agentic framework.

Let's talk about your harness and how to create a strong AI SDLC

Contact Us

Stay connected

with Abstracta

News, articles, and resources on building better software.

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Read about our privacy policy.

Illustration of two people connected by a bridge, one with a laptop and one with a tablet, representing collaboration and bridging communication. End of article about AI SDLC.