Skip to main content
Operating model for an autonomous engineering agent: on the left Project Intent & Requirements, in the middle the agent with the context building blocks Domain Toolchain, Domain Practices and Org Capabilities as well as the cross-project operating model made up of Engineering Method, Engineering Platform and Governance & Conventions, on the right Results and Milestones, at the bottom a control loop back to the intent.
AI

AI agents in engineering need an operating model

An AI agent took my DIY robot to a buildable design. Lesson: tech stack and rules aren't enough. Quality needs an operating model: method, platform, governance.

Automatically translated from German · Read the original

Julian Weyer
Julian Weyer October 1, 2026 · 7 min read

My operating model for an autonomous engineering agent. Intent goes in, milestones come out, and the control loop feeds experience back.

AI ·AI ·Agentic AI ·7 min read

When an AI agent is supposed to develop a product, people usually look at the model and the tools first. Which LLM, which CAD, which simulation? After a few weeks of the robot challenge, I’m convinced that the quality of the result depends on the operating model the agent works in. In other words, on how work is done, decisions are made, results are checked and lessons are learned.

As a reminder: in the Next challenge I set out to have an AI agent develop a small DIY robot autonomously. From requirements through architecture, CAD and firmware to simulation. I define the framework and the requirements; the assembling is up to me. A first design of the robot is now done and simulated (not built yet). What’s more interesting is what I learned along the way.

Following the process and still off target

My first draft of the framework consisted essentially of basic rules and a tech stack. The result was pretty mediocre.

The most striking example: the agent had picked a ready-made kit as the base on its own, a round acrylic chassis with two decks. In its CAD model, there was then a rectangular aluminum frame. The product photo had been sitting in the project folder the whole time. I only noticed it myself, just by looking.

Same kit, two CAD models: on the left the product photo of the round acrylic chassis with two decks, in the middle the agent's CAD model from round 1 with a rectangular plate, on the right the corrected CAD model with round decks.
Left: the kit. Middle: the agent's CAD model from round 1. Right: the state after the correction, thanks to an improved operating model. Since then, a comparison image is mandatory for every purchased part. Photo: roboter-bausatz.de · Chassis model on the right: “Chasis Circular Robot” by Tecneu, GrabCAD

What’s remarkable: the agent had followed all the process rules. It derived requirements, documented decisions, researched with sources. It simply lacked a yardstick for what good engineering work looks like. To put it bluntly, the aluminum frame was a “should be fine” for it. And somehow I had taken it for granted that it would of course be a given for the model to match what is supposed to come out in the end.

The second lesson was about me. From the start, I knew I wanted to separate the rulebook (how does the agent work?) from the project (what is it building?) so that I could reuse it. I hadn’t written that down as a requirement either. It became clear to me when I looked at the first framework drafts: no, we’re not there yet.

Implicit expectations either become explicit requirements, or they don’t get met.

My operating model takes shape

After two or three rounds of revision, this became my operating model in its current form.

On the left, the intent goes in: goals, requirements, constraints, budget. On the right come build packages, prototypes and test reports. At the bottom, a control loop runs back with findings, lessons learned and change requests.

Operating model for an autonomous engineering agent: on the left Project Intent & Requirements, in the middle the agent with the context building blocks Domain Toolchain, Domain Practices and Org Capabilities as well as the cross-project operating model made up of Engineering Method, Engineering Platform and Governance & Conventions, on the right Results and Milestones, at the bottom a control loop back to the intent.
My operating model for an autonomous engineering agent. Intent goes in, milestones come out, and the control loop feeds experience back.

The agent itself consists of several layers. The upper part (blue) depends on the domain and the organization (or in my case: on me): tools such as SysML, CAD, ECAD and simulation, plus design rules, lessons learned and the question of what the organization (=me) can actually manufacture and test. In my “one-man company”, that means: soldering yes, SMD no, no 3D printer, one multimeter. The agent has to plan with what I can actually deliver later on.

The lower part (yellow) is the fixed core of the operating model. It applies across projects and consists of three layers:

Governance & Conventions. Who decides what; how changes are handled; when the agent hands over to the human; how assets are named.

Engineering Platform. Git, containers, skills and automated checks. It could just as well be a full-blown PLM landscape such as Windchill or Teamcenter.

Engineering Method. Model-based and architecture first. Requirements and architecture come first (in my case in SysML v2), then design and code. Every requirement is traceable down to the test case.

If you’re at home in PLM, you won’t find much that’s new here. Development handbook, change management, release process, design guidelines: human teams have had all of this for decades. An agent needs the same, only much more explicit and, in many places, enforced by machines.

What was still missing at the start

Method, platform and ground rules for decisions were there from the start. What was missing was a yardstick for good engineering work. The framework got it in the second round, in the form of design guidelines for the individual disciplines, just as every development department knows them. The lesson from the chassis is one of them.

It was just as important to me that these guidelines don’t stay frozen at day one. The agent records what it has learned and proposes new rules based on that. I decide which of them apply.

When I trust the agent

My article “AI in engineering needs rules – and trust” was about the question of when you can trust an AI’s work. That kind of trustworthiness is exactly what the operating model is meant to establish. I want to be able to rely on the result without checking every line myself.

That requires traceability first: every decision is justified, every requirement is traceable to its test. Then clear responsibilities. The agent settles technical questions on its own; goals, budget, safety and the ground rules themselves it settles only with me. And the platform enforces these boundaries, since a request in the prompt would be too weak over long sessions. Fun fact: when I once changed the ground rules directly myself, without a change request, the agent noticed at the next start and traced it back to me. Caught!

Agentic engineering is not vibe engineering

Andrej Karpathy described vibe coding roughly like this: you give in to the vibe and see what comes out. And the terms vibe engineering and agentic engineering often get mixed up. But what’s happening here is anything but “vibe”. There is a clear framework that is supposed to fulfill requirements against an intent, and to do so traceably.

Building this framework took a lot of brainpower and very little vibe.

In the process I notice something I know from long experience with AI. For focused individual tasks, it works superbly. For larger undertakings, where perhaps only the intent is known, the direction is still very unclear, and even the definition of done can only be put into words vaguely, it only works as a dialogue between human and AI. Without AI, this project wouldn’t exist. But neither would it without me.

And the robot?

What’s taking shape is a shy home robot. It gets startled by light and noise and hides in the dark, preferably under our sofa. It is simulated in a digital replica of our living room, running the same firmware that will later run on the microcontroller.

Under the hood of “proto-1” there are around 50 requirements, all traceable to the test case, 14 documented decisions, an automatically checked SysML v2 architecture and a bill of materials of just under 75 euros (against a set budget of 100 euros). The agent built three bugs into its firmware itself. It also found them itself, in scenario tests and simulation.

Startle simulation in the digital living room, from three camera angles. The robot model here is still the state before the chassis correction.
Floor plan of the living room with eight driving paths of the robot from the simulation study on searching for darkness.
Dark-search study: eight starting points, eight driving paths. The only truly dark hiding spot is under the sofa.

Next up is the build. Then we’ll see whether the design holds up and how much the simulation was worth. According to the requirements, the little one flees from light and noise. Chasing me isn’t in there anywhere. I’m curious whether it will stick to that 😉

A gentle hint: if you’re wondering what an operating model for AI agents in your own engineering could look like, with requirements, change process and Digital Thread, my colleagues at BHC and PROSTEP and I are happy to help.

Questions and answers

What is an operating model for AI agents in engineering?

The framework an agent works in, independent of any single project: an engineering method (in my case model-based, architecture first), an engineering platform (Git, containers, skills, automated checks) and governance with conventions (decision rights, change process, hand-overs to the human, naming). On top of that comes a control loop that turns experience from the project into new rules.

How does agentic engineering differ from vibe coding?

With vibe coding, you roughly describe what you want and let the result surprise you. Agentic engineering works against an explicit intent: requirements with IDs, documented decisions, change requests and checks, so that every result can be traced back to its requirement.