Yes, it is safe to run AI-generated PowerShell in production, but only with a sign-off gate in front of it. A script an assistant wrote is no more dangerous than one an engineer wrote at four o'clock on a Friday. What makes it risky is that most organisations have no rule covering it. There is a change process for a firewall rule and a release process for an application, and then there is a script that arrived in a chat window and ran with a Global Administrator token. The missing control is not a model setting. It is a named approver, a defined list of what that person checks, and a record of what ran and why.

We hold ISO 27001, ISO 9001, ISO/IEC 42001 and Cyber Essentials, so this comes from being asked these questions in assessments.

The failure is the missing gate, not the model

Most of the argument about AI in operations is really an argument about model quality, which is the wrong argument. The same organisation already runs scripts from contractors, from forums, and from staff who left in 2019. Provenance was never the control. Review was.

What genuinely changes is throughput. One engineer can now produce twenty plausible remediation plans in an afternoon, so the constraint moves from generation to review, and review is the part nobody staffed. A plan that reads confidently and cites the right cmdlets slips past a tired reviewer, because nothing in it is obviously wrong. What is missing is a bounded scope, and absence is harder to spot than error.

So the first question is not "can we trust the AI" but "what does the person clicking approve actually check".

Flow diagram of a sign-off gate for AI-generated operational change. A request arrives from an engineer or a service desk ticket. The AI assistant drafts a plan and a script, recording the request and the tool version. The draft passes into a review gate where a named reviewer checks seven things: scope and filters, destructive operations, dry-run support, idempotency, blast radius, the identity the work will run under, and the rollback path. Three outcomes leave the gate: approved and executed, returned for amendment, or rejected. Approved work runs under a named service identity and writes an execution record holding who approved it, what ran, what changed and the result. An arrow returns from that record to the gate, labelled promotion review, because a workflow that passes repeatedly is what earns a wider autonomy lane.

The gate is the product: everything before it is a draft, and everything after it is evidence.

Three lanes of autonomy

Whether AI is "allowed to act" is the wrong question, because the answer differs per task. Sort work into three lanes, granting the lane to the workflow rather than the tool.

Read-only. Log analysis, drafting documentation, explaining what a policy does. Nothing changes state, so the failure mode is a wrong answer rather than a wrong action. Everything starts here.

Propose and approve. The assistant produces the plan and the script; a human reads it and executes it. Most useful operational work belongs here, and this is where the gate lives. Approval has to be an act with a record. If the process is "the engineer looks at it and runs it", that is the third lane with extra steps.

Execute within policy. The assistant runs the change inside a pre-authorised boundary: a defined set of operations, a defined target scope, a blast radius ceiling and automatic stop conditions. Good candidates are narrow, repetitive, reversible and monitored, such as applying a documented licence reclaim to accounts already confirmed as leavers.

A workflow earns the third lane once it has run in the second enough times to be boring: approved repeatedly without amendment, deterministic, bounded structurally rather than by a typed filter, and reversibly tested. Demote on any surprise, and never transfer a lane to a similar-looking task.

What a reviewer actually checks

This is the list missing from most policies. Put it next to the approve button and make the reviewer record which checks they made.

Scope and filters. What does this touch, and what defines that set? A typed filter is a promise; a reviewed group membership is a boundary. The classic incident is a filter that silently matches everything because a property was null on more objects than expected.

Destructive operations. Look for remove, delete, disable, reset, and any write to a production configuration. Reads are cheap to be wrong about; writes are not.

Dry-run support. Does it support a preview, and has the preview been run and read? In PowerShell that means real ShouldProcess handling so -WhatIf works, not a comment claiming the script is safe.

# Unbounded: matches every disabled account in the tenant, no preview, no confirmation
Get-MgUser -Filter 'accountEnabled eq false' -All | Remove-MgUser
# Bounded: a reviewed membership defines the scope, and the preview runs first
$leavers = Get-MgGroupMember -GroupId $ConfirmedLeaversGroupId -All
$leavers | ForEach-Object { Remove-MgUser -UserId $_.Id -WhatIf }

Idempotency. If this runs twice, does the second run do nothing or do damage? Ask what a retry after a partial failure does, because that is the common case.

Blast radius. Not "is this correct" but "if it is wrong, how many people notice and how fast". One licence assignment and one conditional access policy give very different answers.

Which identity it runs under. The check people skip. If the work executes inside an engineer's privileged session, every log entry says a human did it. Unattended work belongs to a workload identity holding only the permissions the task needs, the same discipline the current question set pushes on service accounts in our Microsoft 365 auto-fail checklist.

Rollback path. What undoes this, and has anyone tested it? Restoring from backup is not a rollback path for a directory object, and what cannot be undone needs a higher approval, not a faster one.

What the standards say, and what they leave to you

Neither ISO 27001 nor Cyber Essentials contains a control reading "you may not run AI-generated scripts". Waiting for one is a mistake, because the obligations already apply and simply do not use the word.

ISO 27001:2022 asks for controlled change, logging that lets you reconstruct events, separation of privileged access, and defined configuration management. An action originating from an assistant is a change and needs the same authorisation and record as any other. ISO/IEC 42001 goes further, asking for a management system around AI use: defined purpose, defined roles, risk assessment and monitoring of behaviour in practice.

Cyber Essentials mostly does not care how a script was written. It cares that administrative access is separated and that the accounts capable of making the change are controlled. The connection is indirect but real: the moment an assistant needs standing privilege to be useful, you have created exactly the permanently elevated identity the scheme is unhappy about.

What they leave to you is the judgement call: none of them will tell you which operations belong in the third lane, and what an assessor asks for is the reasoning and the record behind that decision. Our mapping of Microsoft 365 settings to Cyber Essentials and ISO 27001 covers the control picture underneath.

The evidence record an assessor accepts

Approval that leaves no trace did not happen. A workable record per executed action holds the trigger, the requester's intent in their words, the full approved artefact rather than a summary, the tool and version that produced it, the named reviewer and which checks they recorded, the decision and timestamp, the executing identity, the target scope resolved at execution time, the result including partial failures, and any rollback taken.

Two properties matter more than completeness. It has to be queryable, because an assessor asks about a date range rather than a folder. And it has to outlast the platform's own logs, which expire sooner than most people assume.

Where it breaks

Shared administrative accounts. The approval record names a person, the execution record names an account, and nothing joins them. Fix the accounts first.

The assistant acting under a human's credentials. Convenient, and it destroys attribution. Every downstream log then shows a person performing actions they may not have read closely.

Log retention. Your approval record may live for years while the platform record of what changed expires in weeks. You prove somebody approved something and fail to prove what it did.

Delegated access at managed-service providers. When work lands in a customer tenant through delegated access, decide which side holds the approval record and how the customer obtains it, alongside what to report every month.

Approval theatre. If the reviewer approves everything within seconds, the gate is decorative. Measure amendment rate, not approval rate.

An adoption sequence with stop conditions

  1. Inventory what already happens. AI-generated scripts are already being run. You need the honest baseline.
  2. Write the reviewer checklist. The seven checks above. Stop if you cannot name who reviews.
  3. Run everything in the propose-and-approve lane for a quarter. Stop if the amendment rate is near zero, because that means nobody is reading.
  4. Give unattended work its own identity. No standing Global Administrator, no shared accounts, scoped permissions per workflow. Nothing after this works without it.
  5. Promote two or three workflows, with stop conditions set in advance: an object count ceiling, an error threshold, an out-of-hours block.
  6. Review quarterly and demote freely.

Allow two quarters. The slow part is never the technology; it is agreeing who owns the decision.

Where EtherAssist fits

That gate is what EtherAssist is built around. It produces the plan and the script with the operational context of a Microsoft estate behind it, holds the work at a human approval step, and writes the execution record on the other side, so approval and evidence are one artefact rather than two systems to reconcile later.

That is what our agentic operations route covers: scoped execution with human approval and an audit trail, distinct from the broader IT operations and compliance work. If the driver is an upcoming assessment, ISO compliance and audit readiness fits better, and what agentic AI means for IT teams sets the vocabulary. EtherAssist is GBP 16.00 per user per month with a 14-day trial, so a real workflow can go through the gate first.

Explore agentic operations to see how scoped execution, human approval, and the evidence record fit together in one workflow.