AI agents can remove many handovers, but only when embedded across the software-defined vehicle lifecycle rather than deployed as isolated coding assistants. The alternative is not another tool. It is a lifecycle. Its limit will not be how much software AI can produce, but how much AI-generated work the organization can responsibly accept. We call this accountability bandwidth.

Start with what the software-defined vehicle has already changed. The car’s software has broken free from the car’s manufacturing calendar. Metal still has a hard milestone at the start of production. Software runs a loop that opens long before that date and keeps turning for as long as the vehicle is on the road: requirements, architecture, code, testing, simulation, validation, release, field feedback, and the next requirement. That loop is the product now, and how fast it turns is the competitive question.

Thus, the interesting question about AI agents in automotive is not whether they can write code. They can. The question is what every turn of that loop looks like when there is an agent at each station – and what has to be true for the result to be admissible in a vehicle that carries people.

One loop, six places to delegate

One clarification before the walk-through, because it changes how the rest reads. What sits at those stations is not one general-purpose assistant asked to be helpful. It is a curated set of agents with narrow jobs and deliberately limited permissions: a requirements tracer, a test author working from the requirement, a coding-standard and architecture conformance reviewer, a variant-difference analyst, a diagnostics test generator, an evidence writer. Each one has its own instructions, its own context and its own restricted set of tools. That is what makes the output reviewable, and it is also what makes the capability transferable from one program to the next instead of living in the habits of whoever set it up.

  1. Requirements and intent: An agent can draft a requirement, standardize it against the existing estate, identify the requirements it contradicts, and trace it to the tests that would prove it. What it cannot do is invent the intent. And here the first hard lesson arrives: a vague requirement is no longer a problem that surfaces slowly in review. It gets executed incorrectly at scale in an afternoon. Specification quality stops being a documentation virtue and becomes a throughput constraint. Most organizations discover, at this exact point, that their requirements were written for human interpretation and that a machine cannot read them at all.
  2. Architecture and contracts: Every component makes promises about timing, ranges, units, error behavior. In most estates those promises live in a document upstream and nowhere in the code itself. An agent can read both sides and reconcile them, and on a real legacy estate that reconciliation fails immediately. This is worth stating precisely because it is the norm rather than the exception: the contract is declared somewhere but invisible where it matters.
  3. Beyond code generation: This is the least interesting station, which tends to surprise people. By volume, most vehicle software is the kind of work where correctness can be demonstrated rather than argued: drivers,communication and data marshalling, diagnostics, state machines, platform glue, porting from one platform generation to the next. Here the human reviews the contract, not the lines. A smaller share is application and comfort logic, where behavior has to be tested rather than proved. A smaller share still – perception, prediction, and planning – is where an agent should build the test apparatus and never the function itself.
  4. Test and virtual validation: The rule that makes the whole construction hold is the oldest control in engineering: whoever does the work does not sign it off. In agentic terms, the agent that writes the test is not the agent that wrote the code, and it derives that test from the requirement rather than from the implementation. This is also where virtualization stops being an adjacent investment and becomes a precondition. Agents deliver value only where feedback cycles are short. An agent that waits two weeks for hardware is an expensive way of producing guesses.
  5. Release and evidence: Traceability is normally reconstructed under pressure, by people, shortly before someone in authority needs to see it. In an agentic loop it is generated as a by-product: which agent produced this artifact, based on which instruction and context, checked by which tool, and accepted by which named person. Assembled continuously, it is nearly free. Reconstructed afterwards, it is one of the most expensive activities in the entire program. Code-to-car becomes spec-to-car.
  6. The fleet, and back to the beginning: Field data is triaged, failures are reproduced in the virtual environment, the update is staged first to the vehicles most likely to expose the problem, and the next requirement is drafted from the results. The loop closes – and the interval between a customer’s experience and an engineering change becomes something you can deliberately shorten.

The bottleneck moves; it does not disappear

Add all of that up and something uncomfortable falls out. Agents compress production and inflate assurance. More gets produced, so more must be reviewed, integrated, evidenced and accepted. An organization that installs agents and keeps its present shape converts agentic productivity into queue rather than speed. By most industry forecasts, integration and validation are already the fastest-growing cost line in vehicle software. Agents relocate that bottleneck into verification; they do not relieve it.

The real capacity limit is accountability bandwidth: the number of independent acceptance decisions that named, competent and legally accountable people can make without assurance quality deteriorating. It is finite and rarely measured. Improve the evidence per decision or reduce the decisions requiring a signature. Both are lifecycle design choices, not tool procurement choices.

Fewer routine tasks, more accountability

The uncomfortable question is no longer which tasks can be automated. Many can. The real question is which capabilities become more valuable, and who remains qualified to review, accept, and sign off the work.

Producing software becomes easier. Assessing whether it is correct, safe, robust, and fit for purpose does not. As more work is generated automatically, the value shifts toward engineers who can challenge assumptions, identify risks, and take responsibility for the outcome. The engineer who never works through the hard version of a problem is less likely to develop the judgment needed to sign off the result with confidence. The productivity gain is immediate. The capability debt emerges over time.

The risk sits where you least expect it

Now the counter-intuitive part, and the most useful sentence in this article: agent-authored software is safest exactly where the safety requirements are strictest. At the highest criticality levels the standards already call for a formidable detection apparatus – structural coverage down to individual decisions, fault injection, resource analysis, formal inspection instead of a walkthrough, confirmation by someone independent of the department that did the work. Point an agent at that station and most of the machinery needed to catch its mistakes is already installed, already budgeted and already trusted.

At the other end – comfort features, convenience functions, everything nobody classifies as safety-relevant – almost none of it is required. Same generator, same classes of mistake, nothing configured to catch them. And the instinct in every organization is to let agents run unsupervised precisely there, because doing so feels harmless. So the practical rule is to configure the factory inversely to intuition, and to choose the first serious pilot at high criticality rather than low. It is the cheaper evidence package and the more persuasive one.

One further governance point, which no standard recognizes yet, so we offer it as our own proposal: two agents running on the same model do not constitute two independent reviewers . They share a training distribution and, therefore, a family of blind spots: one reviewer with amnesia rather than two people. Human error is largely independent; agent error is correlated. Wherever independence is claimed, model diversity deserves to be governed the way dual-sourcing is governed. Achieving that diversity organizationally may require restructuring. Here, it may require only one line of configuration.

What does not move

Delegating more must not mean controlling less. Hazard analysis, safety goals, criticality classification, residual-risk acceptance and the final signature stay with named, accountable people. AI agents cannot carry accountability for the vehicle. Agents build it. Humans sign it.

Three numbers before you buy anything

None of this starts with an agent. It starts with a baseline, and three numbers are enough to establish one.

  1. What share of your requirements can a machine read without prior human interpretation?
  2. How long does it take, from making a change, to producing complete, traceable evidence for it?
  3. And in one value stream, how many times does work stop so that a person can carry information from one system into another? That third number is your real cycle time, and it is usually the one that silences the room.

Then take a single slice of the loop, instrument it, and design backward from the number you actually want to move – cost, cycle time or ramp-up time. Fund virtualization properly, because it determines whether agents deliver value at all. Separate authorship from acceptance before you scale, not after. And pick the hard pilot: the loop you least want to touch is the one that will teach you the most about whether any of this is real in your organization.

We will be at the Mondial de l’Auto in Paris from October 12th to 18th, with a full conference day on October 14th, showing this loop end to end rather than one station of it. Bring the value stream you would most like to shorten.