Humanist plum-toned illustration of a person holding a key at a garden threshold while cobalt context lines stop before the gate.

Which Decisions Should Never Be Delegated to an AI Agent

By Tim Crossley · 2026-06-26 · 9 min read
Share
The short version

AI agents can prepare decisions, gather context, draft recommendations, and route work to the right person. The line is final authority: decisions that change promises, policy, money, company knowledge, or accountability should stay with people.

On this page

An AI agent is a system with a job. They can watch for work, gather context, run a skill, update a record, draft a reply, flag an exception, and bring the result back to the right person. In a business setting, that can remove a great deal of repeated preparation from the day. Preparation, though, is not the same as authority.

Delegating a decision means the agent is allowed to make the final call and carry out the consequence without a person approving it. That is the line to treat carefully. An agent can prepare the work around many decisions, including important ones. They can assemble the facts, apply the known standard, show the likely answer, and explain what would happen next. But when the decision changes what the company promises, what it believes, what it spends, who it trusts, or what risk it accepts, the final authority should stay with a person.

This is not because people are always wiser in the moment. People miss things. They get tired. They route hard calls through memory and habit. The reason to keep certain decisions human is that the business needs someone accountable for the consequence. A system can carry process. A person has to carry responsibility.

The useful question is not "Can AI make this decision?" Sometimes, technically, it can. The better question is, "If this decision goes wrong, who understands the tradeoff well enough to own it?"

The Agent Can Prepare More Than They Can Decide

A lot of work that feels like decision-making is actually preparation around a decision.

A lead comes in, and the business needs to know whether it fits. A proposal needs a scope recommendation. A project looks off track, but someone needs to know whether it is truly at risk or just noisy. A customer asks for an exception. A manager needs the weekly numbers with the unusual patterns pulled forward.

An agent can help with all of that by reading the inquiry, checking the customer's history, comparing the request against the company's criteria, gathering the missing facts, drafting the likely response, and holding the next step for review. In many companies, that preparation is the expensive part. The final decision may take five minutes. Rebuilding the background takes twenty. That is a good place for AI.

The mistake is letting the usefulness of preparation blur into the ownership of judgment. If an agent says, "This lead looks like a strong fit because the project type matches our current offer, the timeline is inside our delivery window, and the referral source is trusted," that is useful. If the agent rejects the lead, changes the CRM, and sends a firm no without review, that may be too much authority unless the company has already defined that exact case, tested it, and accepted the risk.

The difference is not subtle inside the business. Prepared work gives a person leverage. Unreviewed judgment can create a new kind of supervision: the team has to watch the system, wonder what it decided, and clean up the places where the decision was technically plausible but commercially wrong.

Do Not Delegate Decisions the Company Has Not Made

An agent should not be asked to settle ambiguity the company has avoided.

This shows up whenever a workflow is slow because the standard is not real yet. Maybe the pricing policy changes depending on who asks. Maybe leadership has never decided whether a certain customer type is still worth pursuing. Maybe the delivery team has an unofficial limit that sales has never accepted. Maybe everyone agrees the process is broken, but no one has chosen the new one.

AI will not solve that kind of uncertainty. It will expose it.

If the company has not decided the policy, the agent should not invent one. If two leaders would answer differently, the agent should not become the tie-breaker. If the team does not agree on what good looks like, the first job is to write the standard, not automate around the absence of one.

The agent's useful role is narrower in that situation: surface the conflict, show that the current context contains two different pricing rules, route the exception to the person who owns the domain, and draft the decision record once the person decides. The business has to make the call before the system can carry it.

This is one of the sharper tests for readiness. If you cannot tell the agent what the rule is, the agent is not ready to apply it.

Do Not Delegate New Promises to Customers

Customer-facing work deserves a special kind of caution because a message becomes a promise the moment it leaves the company.

The preparatory work can still be substantial. An agent can draft replies, summarize the customer's history, pull in approved language, identify whether a request matches the company's standard offer, prepare a response that is probably right, and hold it for approval.

Creating a new commitment is different.

The dangerous moments are often small. "We can make that date work." "We can include that at no extra charge." "We should be able to handle the rush." "This is covered." "That will not be a problem." Those sentences may read like ordinary service language, but each one changes the shape of the relationship. A delivery team may inherit the promise. Finance may inherit the exception. Leadership may inherit the risk.

If the message is routine and the company has already approved the exact lane, the agent may eventually earn room. An order-status reply with no policy choice inside it is different from a custom promise. A scheduling confirmation is different from changing the scope. A receipt is different from approving a refund.

This is where earned autonomy matters. Autonomy should attach to a specific kind of action, not to the agent as a whole. The agent may earn the right to send one category of routine reply while still holding every pricing exception, refund, scope change, or delicate customer message for review.

The line is not "customer-facing equals impossible." The line is whether the customer-facing action contains a new promise, an exception, a judgment call, or a risk the company has not already bounded.

Do Not Delegate Exceptions That Change Money or Risk

Money decisions are often disguised as small operational decisions.

A discount. A waived fee. A refund. A payment-term exception. A rush request. A nonstandard scope item. A customer who wants the company to start before paperwork is complete. A vendor who needs approval outside the normal range. None of those has to be dramatic to matter.

An agent can prepare the case by showing the customer's history, the standard policy, the current margin, the open invoices, the project status, and the likely consequence of saying yes or no. The internal recommendation can be drafted and routed to the right person.

But the final approval should stay with someone who understands the commercial context.

The reason is not that an agent cannot read a policy. The reason is that exceptions create precedent. A one-time discount can teach a customer how to negotiate. A waived fee can become an expectation. A rushed timeline can train the team to absorb chaos. A small approval can carry a margin consequence that only makes sense when someone sees the broader account, capacity, and strategic value.

This is one reason a narrow access boundary is not enough by itself. The agent may have permission to read the right information and still not have authority to decide what the business is willing to risk. Access answers "What may the agent see?" Authority answers "What may the agent decide?" They are related, but they are not the same question.

Do Not Delegate People Decisions

People decisions are rarely clean inputs with clean outputs.

Hiring, firing, compensation, role changes, performance concerns, conflict between teammates, sensitive feedback, and questions of trust all involve facts that are incomplete, social context that may not be written down, and consequences that affect real people inside the company.

Around those decisions, an agent can still be useful. Interview notes can be summarized. Feedback can be organized. A role description can be compared to a candidate's experience. A manager can get help drafting a clearer internal note. The team can be reminded which policy applies.

They should not decide who is hired, who is disciplined, who is trusted, who is promoted, or who is no longer part of the company.

The issue is not only accuracy. It is legitimacy. People inside a company need to know that consequential calls about them are made by accountable people, not by software applying a pattern to an incomplete record. Even when a person uses AI to prepare, the person has to remain visible as the decision-maker.

This matters beyond formal employment decisions. A small team often runs on trust, memory, and context that never appears cleanly in the tools. Who can handle a sensitive customer? Who needs support rather than correction? Who is stretched too thin? Who is ready for more responsibility? Those calls belong with people who understand the whole situation and are willing to be answerable for it.

Do Not Let Agents Rewrite the Company's Official Memory

The company record is another place where final authority matters.

An agent may notice that a policy is outdated. They may see that a customer preference keeps recurring. They may catch that a proposal template still includes old language. They may identify that three corrections all point to the same missing rule.

That kind of notice is valuable, exactly the kind of learning a good system should surface. But an agent should not quietly rewrite the company's official knowledge. In Protobase terms, durable company knowledge changes through proposal and review. The system proposes. People approve. That loop can feel slower than letting software update the record on its own, but it is the reason the record can be trusted later.

This is especially important because official memory gets reused. It shapes future drafts, future recommendations, future onboarding, future reports, and future agent behavior. A bad reply is one mistake. A bad rule in the company's official context can become many mistakes.

The agent can write the proposed change, explain why the change seems needed, include sources and examples, and route the proposal to the person who owns that area of the business. The person has to decide whether the company should remember it as true.

That distinction keeps learning alive without letting memory become loose.

Some Decisions Can Earn More Room

The word "never" needs care. Some actions that start held for review can later run automatically. A routine reply that was approved unchanged for months may not need the same review forever. A standard status update may be safe once the inputs are reliable. A record cleanup task may become automatic after the system proves it can identify duplicates correctly and undo mistakes cleanly. But that freedom should be narrow, earned, and reversible.

The agent does not become broadly trusted because one responsibility worked well. A support agent that earns the right to send order-status replies has not earned the right to approve refunds. A reporting agent that prepares reliable weekly summaries has not earned the right to change the company's targets. Trust belongs to the specific action, inside specific limits, with a record behind it.

This is where an eval becomes more than a technical detail. The business needs a way to test whether the agent followed the standard, used the right context, stayed inside the access boundary, and produced work that people approve without constant correction. If the work drifts, the autonomy should shrink. If the context changes, the action may need to go back to review.

Good autonomy is not a switch. It is a record of reliable work attached to one narrow responsibility.

A Practical Test

When deciding whether an agent should own an action, ask what would have to be true for the company to be comfortable with the result happening without review.

The action needs a clear rule. The inputs need to be available. The downside needs to be understood. The output needs to be inspectable after the fact. A mistaken action needs a clean recovery path. The company needs a person who owns the policy behind it. And the agent needs a record showing that this exact kind of work has been reliable.

If those conditions are missing, the agent can still prepare the work: gather context, draft the likely action, show the reasoning, and hold the step for approval. That is often enough to create real leverage without pretending the system owns more judgment than it does.

The practical boundary is simple to say, even when the implementation takes care:

Let agents prepare work where preparation creates leverage. Let them act automatically where the action is narrow, tested, reversible, and already bounded. Keep people responsible where the decision changes promises, policy, money, trust, official knowledge, or risk.

The Deeper Reason

AI can make a company faster at producing answers. That is useful only if the company also knows which answers deserve authority.

In an established business, many expensive decisions are expensive because they carry the company's judgment in a concentrated form. What will we promise? What risk will we accept? What standard do we hold? What do we remember as true? Who is accountable when this turns out to have consequences?

Those questions do not disappear because an agent can draft a plausible answer.

The better system is not one where people approve every small thing forever. That would make AI timid and tedious. The better system is one where the business separates preparation from authority, gives agents narrow jobs, measures the work, expands only where the record supports it, and keeps consequential judgment attached to accountable people.

Agents become useful without becoming slippery when they reduce the repeated work around decisions while leaving the decision itself where it belongs. The routine can move. The judgment needs a name on it.

Questions this note answers

A few direct answers.

Which decisions should AI agents never make on their own?

AI agents should not make final decisions that change company policy, alter customer promises, approve unusual financial terms, resolve sensitive people issues, rewrite official company knowledge, or accept risk the business has not already defined.

Can an AI agent help with decisions it should not own?

Yes. An agent can gather context, compare the situation to approved standards, draft options, flag missing information, and route the work to the right person. Preparing a decision is different from owning the decision.

When can an AI agent act automatically?

An agent can act automatically only for a narrow action that has clear limits, low downside, reliable tests, review history, and a way to pull the freedom back if the work drifts or context changes.

Get in touch

Let’s talk.

If you run an established business with a team and too much of it still routes through you, choose a time below. Thirty minutes on Google Meet, at no cost. Tell us how the business runs and where work gets stuck, and we’ll tell you honestly whether Protobase is a good fit.