Skip to main content
Trust fall: a woman with her arms crossed lets herself fall backwards and is caught by a humanoid robot; the word “TRUST” in the background.
AI

AI in engineering needs rules – and trust

Neither blind trust nor total control: how companies can delegate AI in engineering properly – task-specific, based on observation and revocable.

Automatically translated from German · Read the original

Julian Weyer
Julian Weyer August 30, 2026 · 15 min read
AI ·AI ·PLM ·15 min read

That AI in PLM and engineering needs rules is by now common knowledge. It’s about data access, responsibilities, approvals, evidence, and the question of which tasks a system may take on at all. And in the end, a human being has to bear the responsibility.

I’d like to add something to this important point: AI in engineering needs not only rules, but also trust.

I don’t mean that sarcastically. Nor is it a plea to simply give an AI an advance of trust. Trust is meant constructively here: as a precondition for AI being able to deliver its value at all. If you can’t entrust an AI with a task and some room to act, you may end up with a well-secured system – but probably nothing gained.

At the same time, trust doesn’t come from a decision. It can’t be decreed in a policy, nor can it be replaced by having a human confirm every single action. Trust has to develop: through active engagement with the AI, through limited experimentation, through observing its results, and through a growing understanding of its capabilities and limits.

The real management question is therefore not just: What rules does the AI need? It is also: What are we prepared to rely on it for, under which conditions, and with which consequences?

This is especially tangible in engineering: when an AI summarizes a requirements meeting, flags a contradiction, or searches for similar components, a mistake is usually(!) just annoying, but manageable. When the same AI changes a released bill of materials, makes a compliance statement, or triggers a change process all the way into production, the situation looks different.

Trust is not a blank check

“Trust is good, control is better” is easily said and sounds right. After all, the saying contains an important core: results shouldn’t be accepted unchecked. When dealing with AI, however, it is misleading if control is equated with checking every single work step.

Full control of every individual step doesn’t scale. Confirming every AI action one by one doesn’t automatically create more safety. As the number of decisions grows, time pressure, habituation, and approval fatigue loom. Research on automation bias describes exactly this effect: people working with an automated system gradually check its outputs less carefully and overlook errors that they would have noticed without the automation. In the end, you may be confirming unseen precisely what you actually meant to check carefully.

So control isn’t free. It costs time, attention, and decision-making capacity. And it can itself become a risk if it is applied in the wrong places.

Blind trust is no alternative, though. A plausibly worded output (keyword: AI slop) is not yet a reliable result. An AI can seem convincing even though it lacks context, data is linked incorrectly, or it is working on a task outside its suitable field of use.

The decisive question is therefore not: Do we trust the AI? But rather: What are we prepared to rely on it for, under which conditions, and with which consequences?

Clarifying the term

Trust means delegating under uncertainty

In trust research, trust is not understood as certainty that the other party will behave correctly. Trust is rather the willingness to engage with another party under uncertainty and despite a degree of vulnerability. Mayer, Davis, and Schoorman described this perspective for organizations in a highly influential way.

For the use of automation, a similar distinction matters: the goal is not as much trust as possible, but appropriate expectations of reliability. Lee and See speak of appropriate reliance in this context: people should rely on automation when it is suited to a task – while remaining able to observe, question, or reject its results.

That sounds abstract, but in engineering it quickly becomes concrete. Trust in an AI is not global. It doesn’t say: “We trust this model.” It says something more like:

We trust this system for this task, with this data, in this process, and within these limits.

At least three questions are decisive here:

Capability. Can the AI perform this specific task well enough?

Reliability. Does it work consistently under comparable conditions and within the intended framework?

Fitness for purpose. Is this system suited to the purpose we are using it for?

The last point in particular is underestimated. A generic AI can be useful for many tasks. It doesn’t follow, however, that it is suited to every task. A system that summarizes texts well is not automatically a system that should change released product data.

An AI has no character in the human sense and no sense of responsibility of its own. But it is not a completely neutral entity either. Training data, model behavior, system instructions, and the context provided shape how it responds, what it prioritizes, what it leaves out, and how it deals with uncertainty. This “attitude” is not personality. It is nevertheless relevant to the question of whether the system fits the company’s purpose and rules.

Trust emerges through active engagement

I like to use the following analogy: With a new employee or a new team, you don’t yet know exactly what to expect at the beginning. You know the CV or the team description. But you don’t know their abilities from your own experience. You don’t know how someone handles unclear requirements, time pressure, or mistakes. That’s why collaboration usually begins with limited tasks, questions, observation, and feedback.

Over time, a picture emerges: What can this person do? Where do they work reliably? When do they need support? What responsibility can they take on? As this picture becomes more stable, detailed control can recede.

The analogy of the new employee also helps when dealing with AI. It has an important limit, though: a suitable employee can understand rules, ask questions, learn from experience, and adapt their behavior to new situations. An AI doesn’t do this automatically in the same way. Its working framework (and the learning loop) has to be established more through technical and organizational means.

That is why trust in AI doesn’t come from a decision, nor from a presentation on the capabilities of the latest model. It comes from active engagement with the specific system:

The control loop that matters for this learning cycle is the following:

try out → observe → classify → limit or expand → observe again

Trust research describes a similar progression. Lewicki and Bunker distinguish calculus-based trust, which rests on safeguards and sanctions, from trust that arises from repeated interaction and observed predictability – and finally from trust based on shared values and identification. The third stage is not available for AI. That is exactly why dealing with AI remains permanently dependent on observed predictability: here, trust doesn’t grow from a single great achievement, but from many experiences that stabilize an expectation – and from conditions that make this observation possible in the first place.

This has an immediate consequence for companies: anyone who doesn’t engage actively with AI can develop neither appropriate trust nor appropriate boundaries. A company that merely waits for a better tool or model misses the actual learning task. A powerful model without sufficient context, reliable data, and suitable checks is just a faster route to a misunderstood result.

AI needs a defined “playing field” – not just an approval

You can look at this playing field from two perspectives. On the field itself, the question is how the AI may work. Seen from the sideline, the question is how the organization deals with its results.

The first perspective could be called the framework for action and guardrails for the AI. Among other things, it includes:

Task pass for AI systems: six dimensions – task, context, permitted action, evidence, responsibility, rollback – plus three examples with increasing autonomy: part search (search & suggest), change impact (prepare & analyze), and bill of materials (write access blocked).
A framework for action can be thought of as a task pass: six fixed questions – and a separate answer to each of them for every use case.

Guardrails are more than notes in the system prompt. A limit that the agent can override or circumvent itself is not a reliable limit. Critical restrictions should therefore be enforced outside the model wherever possible – for example through access rights, fixed approval thresholds, API gateways, or process logic.

The second side is the organization’s oversight and accountability architecture. Here it must be clear:

Governance and guardrails are therefore not the same thing. Governance defines the organizational framework: What is permitted, who is responsible, and how is it monitored? Guardrails translate parts of this framework into technical restrictions at runtime.

The permission architecture as a four-level funnel: governance (purpose, responsibility, escalation) steers guardrails (rights, thresholds, API limits, logs), which in turn frame the working context (data, model, rules, tools) and the task (analysis, suggestion, action). Connected on the right: PLM and production via soft boundaries, ERP via a hard boundary without API access.
Governance says what is permitted. Guardrails enforce it — all the way down to the individual task and its access to PLM, ERP, and production.

The art lies in not applying the same control regime to every task. A simple summarization service shouldn’t go through the same approval process as a system that writes into safety-critical product data. For this, the MIT Center for Information Systems Research (van der Meulen, Jewer, Levallet) proposes the term minimum viable governance: as much governance as is needed to limit risk effectively. Beyond a certain ceiling, the research group argues, governance slows things down more than it protects – decisions pile up, and the real opportunity passes.

This is not meant as a plea for less governance, but certainly for the right governance.

Responsibility requires understanding – but not every technical detail

Tiered governance also means that not every AI action ends up on an inspection table. And that immediately raises an objection. If those responsible don’t have to follow every work step of the AI, how can they take responsibility?

The answer lies in a more precise distinction of the term “understanding.” For responsible oversight, at least four levels need to be told apart:

The first three levels are regularly indispensable for responsible decisions. The fourth can be relevant in special cases, but it is not the same as responsibility and, for complex processes, is only attainable to a limited extent anyway.

The EU AI Act, too, doesn’t require that every internal model computation be fully reconstructable for human oversight of high-risk AI systems. Article 14 focuses on something else: that the people overseeing the system appropriately understand its capabilities and limitations, are consciously alert to automation bias, and can interpret, disregard, or override results. Paragraph 3 also explicitly requires that oversight be matched to the risk, the level of autonomy, and the context of use – exactly the proportionality at issue here.

“Technical detail understanding isn’t always necessary” must therefore not mean that domain understanding is dispensable. A person who may only click an approval button but cannot assess the statement, how it came about, and the possible harm is not exercising effective oversight. They are merely formally part of the process.

The problem is often described as rubber-stamping: the human nominally stays in the loop but in practice becomes a yes-man. Systemic control is therefore only defensible if it is combined with genuine understanding of the process and the risk. The EU AI Act holds deployers accountable for this – Art. 26(2) requires that oversight be assigned to people who have the necessary competence, training, and authority.

Engineering has always controlled hand-offs, not every step of thought

You therefore don’t have to bring process and risk understanding to maximum depth everywhere, but at the decisive points. Engineering has always made this selection deliberately: good development processes don’t control every single activity of a design engineer, a project team, or a supplier (even though that does happen, often for economic and controlling reasons). They set control points where risk, responsibility, or the status of a result changes.

Design reviews, verification and validation, FMEA, approvals, and quality gates serve exactly this purpose. They don’t examine every consideration that led to a result (exception: particularly strictly regulated industries). They examine whether the result meets the requirements, whether relevant risks have been considered, and whether the necessary evidence is in place.

Here, too, the analogy again: organization and control theory distinguishes between behavior control and outcome control. The distinction goes back to Ouchi; Eisenhardt links it to agency theory. Behavior control makes sense when the controller understands the process that turns an input into an output. Outcome control makes sense when results can be clearly described and evaluated.

With a large AI model, complete control of the internal processing is usually not a realistic basis for responsibility. That doesn’t mean nothing can be controlled. It means that control has to be shifted elsewhere: to context, data, rules, evidence, result quality, and hand-offs.

PLM governance for AI therefore doesn’t have to start from zero. It can build on existing lifecycle, approval, and change logic. At the same time, AI forces us to examine these processes for their actual risk impact. A change process that is already slow and bureaucratic doesn’t automatically get better through additional individual approvals.

The level of autonomy must correlate with risk

If control is to start at relevant hand-offs, the organization must decide which hand-offs need how much control. This requires a criterion that goes beyond the mere question “AI or no AI?”

Methods for this have long existed: FMEA, for example, rates a failure risk using three variables: severity (how serious is the impact of the failure?), occurrence (how likely is it?), and detection (how likely is it to be noticed beforehand?). The older approach multiplied these values into the risk priority number (RPN); the harmonized AIAG-VDA FMEA of 2019 replaces it with an action priority. Functional safety works with a related but differently structured set: ISO 26262 derives the required safety level (ASIL) from severity, exposure of the operating situation, and controllability. Controllability is particularly revealing for the question of delegation – can the situation still be recovered once the failure occurs? I’m borrowing it as a fourth question alongside the three FMEA variables.

These variables (the three from FMEA, supplemented by controllability from functional safety) are suitable as a criterion for how much autonomy an AI task can bear.

Severity

How far does an error reach into product, process, and organization – does it stay within the design, or does it continue into procurement, production, approval, or customer communication?

Probability of occurrence

Becomes the question of context quality and result stability: Does the system have the right information – and does it deliver reproducible results with it?

Probability of detection

Is there a control loop (human or machine) that can detect errors, and if so, how reliably? A crack in a weld is a crack in a weld. A wrong paragraph in an AI-generated text reads just like a right one.

Controllability

Largely becomes the question of reversibility: Can an AI-induced action be fully rolled back, in all the systems it has already propagated to?

For the combination of propagation and reversibility, the practitioner discussion on AI agents also uses the term blast radius (borrowed from IT and cloud operations, where it is itself already a metaphor from explosion-effect analysis): the extent of the damage a wrong decision can cause before someone catches it. The catchy rule of thumb that goes with it – oversight must grow with the blast radius – comes from the blog of a vendor of tools for coding agents. The term doesn’t replace the FMEA logic, but it can complement it.

Autonomy should therefore not be set across the board for a model or an agent. It is a deliberate design decision, not an automatic consequence of rising model capabilities. This runs through several autonomy frameworks, from the levels of autonomy of the Knight First Amendment Institute to the autonomy level model of the Cloud Security Alliance. For one task, AI can make suggestions, for another prepare analyses, and for a third act independently within hard limits.

A pragmatic gradation could look like this:

LevelNameWhat the AI may do
01SuggestThe AI produces drafts, summaries, classifications, or search results. What happens with them next remains outside the AI and is decided by humans.
02Prepare and informThe AI links information, analyzes impacts, or populates change data. It influences a decision but does not trigger it on its own.
03Execute within hard limitsThe AI may trigger an action itself if scope, data access, thresholds, logging, rollback, and escalation are clearly regulated.

These levels are not a rigid standard. Depending on the company and use case, more nuanced levels may make sense.

Make AI a leadership priority – but not a bureaucracy

Such criteria can’t be decreed by policy. If you pour them into a form, you get a form and no better judgment about which task can bear how much autonomy. That judgment emerges within the organization, and for that, some guardrails at leadership level are needed.

First: AI must become a shared learning task. It is not just an IT, data, or compliance topic. Engineering, business units, quality, IT, and leadership have to develop a shared understanding of what the system can do, where it fails, and what the consequences of an error would be. This takes hands-on work with real tasks. Low-risk applications are not just productivity tools here, but learning grounds. They help calibrate expectations, recognize error patterns, and base trust on solid experience.

Second: experimenting must be allowed. A company that treats every trial like a production-critical rollout will learn only slowly – or (possibly) worse, will end up going around the official governance. Too much and unsuitable governance can encourage shadow AI: employees look for the faster, unofficial route and thereby withdraw their use from exactly the intended control. Experimenting does not mean ignoring risks, however. It means limiting experiments so that the organization can learn without causing uncontrollable consequences.

Third: autonomy must be tied to evidence. The general feeling “we trust the system by now” should not decide on a higher autonomy level. What counts is concrete evidence: For which task are the results reliable? Under which conditions? Which errors are detected? How quickly can one intervene or roll back? Who decides on promotion and demotion?

Fourth: existing engineering processes should not simply be extended, but reviewed or even rethought. The introduction of AI is an occasion to ask whether existing gates actually address the relevant risks. If you put an additional approval for every AI action on top of an already overloaded change process, you only increase bureaucracy – not safety.

Conclusion

Trust is the result of good conditions

AI will only be able to fully deliver its value in engineering when companies give it not only tasks but also appropriate room to act. This room to act must neither arise from euphoria nor be blocked out of fear.

And: trust in AI is not an on/off switch. It is task-specific, based on observation, and revocable. It grows when a company engages actively with the AI, gets to know its capabilities and limits, and draws conclusions from its results.

This takes good context, clear guardrails, relevant control points, visible results, clarified responsibility, and a functioning way to step autonomy back down.

Those who only control without learning stay trapped in detailed control. Those who only trust without observing mistake hope for governance.

The future of engineering therefore lies neither in complete control of AI nor in its complete autonomy. It lies in the ability to give it room to act precisely where people have understood its performance, its limits, and the consequences of an error.

Questions and answers

Why does AI in engineering need not only rules but also trust?

Because an AI only delivers its full value once it is entrusted with a task and some room to act. Anyone who has every single action approved ends up with a well-secured system – but has gained hardly anything. Trust here is not an advance, but the well-founded willingness to rely on the system for a specific task under uncertainty.

What is the difference between AI governance and guardrails?

Governance is the organizational framework: What is permitted, who is professionally responsible, who approves, how is it monitored? Guardrails translate parts of it into technical restrictions at runtime – access rights, thresholds, API gateways, process logic. The key point: a limit that the agent can override itself is not a limit. Critical restrictions belong outside the model, enforced there.

How much autonomy should an AI get in PLM and engineering?

As much as the individual task can bear – autonomy is assigned per task, not across the board per model. Three levels provide orientation: Suggest (the AI produces drafts, humans decide how they are used), Prepare and inform (the AI links information and analyzes impacts but doesn’t trigger the decision), and Execute within hard limits (the AI acts itself when scope, data access, logging, rollback, and escalation are regulated).

How do you assess the risk of an AI task?

Through four questions from established engineering methods. From FMEA: propagation (how far does an error reach into product, process, and organization?), context quality (does the system have the right information?), and observability (is the error noticed before it takes effect?). Added to that is controllability from functional safety under ISO 26262, here as reversibility: Can the action be rolled back in all affected systems? Propagation and reversibility together make up the blast radius (blast radius) – and oversight must grow with it.

Is human-in-the-loop enough as oversight of AI systems?

No. A human in the process is not yet effective oversight. For that, this person must receive the relevant information, be able to interpret the result in terms of content, assess its consequences, and actually be allowed to intervene. If time, competence, or the ability to roll back is missing, rubber-stamping results – the human nominally stays in the loop but becomes a yes-man. Complete individual control doesn’t help against this: it creates approval fatigue and thus exactly the automation bias it is supposed to prevent.

What does the EU AI Act require for human oversight of AI?

For high-risk AI systems, Art. 14 does not require that every internal model computation be reconstructable. What it requires is that the people overseeing the system appropriately understand its capabilities and limitations, are consciously alert to automation bias, and can interpret, disregard, or override results. Paragraph 3 requires oversight to be matched to the risk, the level of autonomy, and the context of use; Art. 26(2) holds deployers accountable for assigning it to people with the necessary competence, training, and authority.

Can too much governance hinder the use of AI?

Yes. The MIT Center for Information Systems Research frames this as minimum viable governance: as much governance as is needed to limit risk effectively. Beyond that limit, it slows things down more than it protects – decisions pile up, the opportunity passes. Unsuitable governance also encourages shadow AI: employees turn to the unofficial route and thereby withdraw their use from exactly the intended control.