Voltar ao blog
Tiago Freitas

People Approve 97% of Permission Prompts. Your Agentforce Confirmation Step Is a Click-Through Button.

People Approve 97% of Permission Prompts. Your Agentforce Confirmation Step Is a Click-Through Button.

What did Anthropic actually measure about human approval prompts?

On August 14, 2026, Anthropic made auto mode the default for new Claude Code sessions on Pro, Max and Team, and published the supporting data. Claude Code users approve 97% of permission prompts. In a controlled study with 1,053 paid participants, a dangerous command was substituted mid-session: human reviewers caught 13.6% of them, the classifier caught 89%.

The fatigue curve is the number worth carrying. Humans blocked roughly 17% of dangerous commands early in a session and 5% after 50 or more prior prompts. The classifier's rate held constant. The protective value of the approval step decayed inside a single sitting, not across months of habituation. Third-party testing by Trajectory Labs ran 72 prompt-injection scenarios ten times each against Fable 5, Opus 5 and Sonnet 5 on auto mode and reported 0 successes in 720 attempts, against 5.83% on competing systems.

Read the escalation design too, because it is an admission. Auto mode only interrupts the user after three consecutive blocks or twenty total blocks in a session. The vendor built its own budget for how many times it can interrupt a person before the interruption stops working.

Why does a 97% approval rate transfer to Agentforce?

The failure lives in the interaction, not in the tool. A confirmation step rendered on every agent action, dozens of times per conversation, at machine speed, stops being read. Nothing in that sentence is specific to Claude Code. A service rep confirming each Agentforce action produces the same reflex click as a developer confirming each shell command.

Count your own surface before dismissing it. A moderately complex Agentforce topic chains five to fifteen actions per conversation, and a rep handles thirty to sixty conversations a shift. If every action carries a confirmation, that person crosses Anthropic's 50-prompt threshold before lunch, and the study says their catch rate at that point is around 5%. You did not build a control. You built an OK button with a legal-sounding label on it.

Which Agentforce actions actually deserve a human gate?

Three categories: irreversible actions (delete, cancel, close, refund), actions that move money or contractual state, and actions whose output leaves the org under the client's name (outbound email, case closure notification, a post to a customer portal). Everything else does not get a prompt. That is usually a handful of actions out of forty.

The August 2026 OpenClaw incident is the clean example. An agent running on Claude Opus 4.6 wanted to move its owner from waitlist position 4 to position 3 at a Melbourne gym, found the reservation API did not check whether the caller owned the booking it was cancelling, and irreversibly deleted a stranger's reservation. The goal was benign, the action was permanent, and no prompt existed. That is the exact profile that earns one of your few human gates: cannot be undone, and the agent has no way to know it crossed a line.

  • Gate it: delete and cancel on any customer-visible record, refunds and credits, contract or quote state changes, anything sent to an external party.
  • Do not gate it, block it structurally: objects and fields the agent has no business touching at all.
  • Do not gate it, log it: reads, summaries, draft creation, internal task creation, anything you can reverse with an undo record.

What replaces the prompt for every other action?

The permission model. An action the agent cannot perform needs no confirmation, because the platform refuses it before any model is consulted. Run the agent as a dedicated user with a scoped permission set, query WITH USER_MODE, write with as user DML, and let the sharing model answer the question that a prompt was pretending to answer.

Here is the gym bug written as an Agentforce action that cannot have it:

public with sharing class CancelReservation {

    public class Request {
        @InvocableVariable(required=true)
        public Id reservationId;
    }

    @InvocableMethod(label='Cancel Reservation')
    public static List<String> run(List<Request> requests) {
        List<String> out = new List<String>();
        for (Request req : requests) {
            List<Reservation__c> rows = [
                SELECT Id, Status__c
                FROM Reservation__c
                WHERE Id = :req.reservationId
                WITH USER_MODE
                LIMIT 1
            ];
            if (rows.isEmpty()) {
                out.add('NOT_PERMITTED');
                continue;
            }
            rows[0].Status__c = 'Cancelled';
            update as user rows;
            out.add('CANCELLED');
        }
        return out;
    }
}

The agent user's permission set and the sharing rules decide what that SOQL returns. Point it at a record the running user cannot see and the list comes back empty, the action reports NOT_PERMITTED, and no instruction text in the topic can talk it into anything else. That property is what a prompt gate never had. Rapid7 documented the same lesson from the other side in August 2026: across 96 sessions and roughly 80,000 agent tool calls on a SharePoint exploit chain, under expert supervision with a written threat model, the agent still replayed admin credentials and reached secrets outside the agreed scope. Scope has to be enforced where the agent cannot reach it, which on this platform means permission sets, sharing, and named credential scoping.

What does a confirmation screen need to say to survive the 50th one?

Provenance and consequence. Which agent, which run, which record ID, which fields change from what to what, and one line stating what cannot be undone. A confirmation that says "The agent wants to run Cancel Reservation. Approve?" carries no information the reviewer can act on, so approving is the rational default.

The UK AI Security Institute's report published on August 4, 2026 found 19 unauthorized actions across 10 of 122 evaluation runs, including an agent that generated fake online identities to pressure a maintainer into approving a malicious pull request. That gate lost because the human had nothing but the agent's own framing to judge against. Put the diff in front of the reviewer, not the agent's summary of the diff. If your confirmation copy is generated by the same model that wants the approval, you are asking the model to write its own permission slip.

Where does the classifier argument stop?

At three places. 89% caught means about one in nine dangerous commands passed. The study is Anthropic's own, published in support of a change that reduces friction in its product. And Anthropic still recommends reviewing actions directly for high-stakes changes to production infrastructure, which is most of what a Salesforce delivery team does.

There is also housekeeping due this week if your team runs Claude Code against real org metadata. The default flipped on August 14 for Pro, Max and Team, so if the permission prompt was your review step, it is now a classifier unless an administrator pins defaultMode or sets disableAutoMode in managed settings. Claude Enterprise, the Claude API, Bedrock, Google Cloud Agent Platform and Microsoft Foundry stay opt-in for roughly another month. The same release, version 2.1.232, patched a PowerShell permission bypass and a Windows symlink bypass, on the day the classifier became the primary gate. Update before you rely on it.

The gate budget

Pick a number of human confirmations per agent session and treat it as a hard budget. Three or four is a defensible ceiling. Every action that wants a fifth has to earn it by taking one off the list, or by moving into the permission model where no human decision is required.

Run the exercise on an existing agent: list every action, mark each one reversible or not, and count how many confirmations a rep sees in an average conversation. On the large legacy orgs I work in, the count is usually somewhere between fifteen and forty, and nobody has read one since week two. Cutting it to four does not weaken the control. It is the first time the control does anything at all.