Voltar ao blog
Tiago Freitas

Salesforce in Claude Runs the Org as You. Your Page Layout Rules Do Not Come Along.

Salesforce in Claude Runs the Org as You. Your Page Layout Rules Do Not Come Along.

What ships in Salesforce in Claude, and what actually executes the writes?

Salesforce in Claude is a plugin carrying 37 prebuilt sales skills, in pilot now with open beta in September 2026. It runs on the Headless 360 Hosted MCP Server, in beta since July 2026, exposing four tools: Discover, Describe, Dispatch, and Dispatch (Read-Only). It requires API version 67.0 or higher and per-user OAuth through an External Client App carrying the mcp_api scope.

The plumbing is not new. Headless 360 shipped at TDX in March 2026 and went GA for Enterprise Edition and above in April. Winter '27 release notes list Hosted MCP Servers as a platform feature on the standard release train, so this is a shipping capability rather than a partnership artifact. What changed on 2026-08-26 is the client: the seller sits in Claude, and the org answers.

One detail carries all the weight. Each request executes as the authenticated individual, so the enforcement point is that user's permission set, and every server-side automation fires on agent-initiated DML. A Salesforce executive stated the boundary as: if you do not own the record, the MCP server does not either. The architecture detail above comes from community analysis rather than Salesforce documentation, so confirm the scope name and beta status in Salesforce Help before you put it in a client deck.

Which of your validation still fires when nobody opens Lightning?

Everything enforced server-side still fires on an agent write: validation rules, record-triggered Flows, Apex triggers, blocking duplicate rules, field-level security, sharing, and fields marked required at the field definition. Everything enforced by the interface does not: required-on-page-layout, LWC and Aura client-side checks, quick action prefills, Path guidance, in-app guidance.

This is the same API-versus-UI split every Salesforce developer already knows from data loads. The difference is who trips it. A data load is run by someone who knows they are bypassing the UI. A seller asking Claude to update a deal has no idea a layer disappeared, and neither does the admin who put a required flag on the layout in 2021 and considered the field governed.

Two specific gaps to inventory in a real org:

  • Required on layout, optional in the schema. Grep your page layouts for required field items, then check each field's actual definition. Anything required only on the layout is now optional for every agent write. The fix is a validation rule or a universally required field, not a layout change.
  • Dependent picklists. The parent/child filtering is a UI behavior. The API will happily accept a combination the picklist would never have offered a human. Test this one in your own org before you trust either answer, because the behavior has varied by object and API version.

Record type page layout assignment, compact layouts and Lightning page visibility rules are presentation. They shape what the model can discover through Describe, not what it can write through Dispatch.

Why does one Claude instruction hit a CPU limit a human never sees?

Writes route through the API, which groups up to 200 records into a single transaction. Synchronous Apex gets 10,000 ms of CPU time per transaction. A trigger and Flow stack costing 60 ms per record is invisible when a rep saves one record in Lightning, and fails at roughly 12,000 ms when 200 records land together.

"Update the close date on every open opportunity in my Q3 pipeline" is one sentence to a seller and a bulk operation to the platform. The record-count governors are usually fine: 10,000 DML rows and 100 SOQL queries per transaction leave room. CPU time is the one that bites, because it is the only limit that scales with how much automation your org accumulated rather than how much data the request touched.

The failure mode is ugly on both sides. Apex CPU time limit exceeded rolls back the transaction, so the seller gets a refusal with no useful cause, and an org with 200-record chunking may commit some chunks and roll back others. Partial success on a pipeline update is worse than a clean failure, because now nobody knows which records moved.

How do you measure your own ceiling before the pilot tells you?

Run one batch update in a sandbox with production-shaped data and read Limits.getCpuTime() around the DML. Divide 10,000 ms by the per-record cost to get the largest batch your automation survives. If that number lands under 200, agent-initiated bulk writes will fail at the API's default chunk size and you know it before a pilot user does.

// Anonymous Apex, sandbox only: this writes.
// Measures what your trigger and Flow stack costs per record.
List<Opportunity> batch = [
    SELECT Id, NextStep
    FROM Opportunity
    WHERE IsClosed = false
    LIMIT 200
];

for (Opportunity o : batch) {
    o.NextStep = 'cpu probe';
}

Long before = Limits.getCpuTime();
update batch;
Long spent = Limits.getCpuTime() - before;
Long perRecord = Math.max(1, spent / batch.size());

System.debug('records:    ' + batch.size());
System.debug('cpu ms:     ' + spent);
System.debug('ms/record:  ' + perRecord);
System.debug('safe batch: ' + (10000 / perRecord));

Run it on each object a skill can write, not just Opportunity. Account and Case usually carry heavier stacks than anyone remembers. Run it twice, because the first execution pays for cold Apex compilation and will read high.

If the safe batch comes back at 40, you have three options and only one of them is fast: cut the per-record cost, move work to async, or cap what a skill is allowed to write in one instruction. Measuring first is what lets you pick before a user picks for you.

What does a defensible pilot permission set look like?

Read-only first, scoped to named objects and fields, with write access added one skill at a time. The Hosted MCP Server exposes Dispatch (Read-Only) as a separate tool, so a read-only pilot is a supported configuration and not a workaround. Assign it to a small named group with viewAllRecords off, so record-level sharing stays the boundary.

<PermissionSet xmlns="http://soap.sforce.com/2006/04/metadata">
    <label>Claude Pilot Read Only</label>
    <objectPermissions>
        <object>Opportunity</object>
        <allowRead>true</allowRead>
        <allowCreate>false</allowCreate>
        <allowEdit>false</allowEdit>
        <allowDelete>false</allowDelete>
        <viewAllRecords>false</viewAllRecords>
        <modifyAllRecords>false</modifyAllRecords>
    </objectPermissions>
    <fieldPermissions>
        <field>Opportunity.Amount</field>
        <readable>true</readable>
        <editable>false</editable>
    </fieldPermissions>
    <hasActivationRequired>false</hasActivationRequired>
</PermissionSet>

The read-only phase is where over-permissive profiles surface. A field that was technically readable but never appeared on any layout was invisible for years. Describe will list it and natural language will ask for it. That is a discovery exercise worth running deliberately in week one instead of hearing about it from a seller who read a compensation field out of curiosity.

Keep the audit trail somewhere the agent cannot write. Event Monitoring, Field History Tracking and Setup Audit Trail sit outside the permission set you just granted. A custom logging object the agent can write to is not an audit trail, it is a file the subject of the audit holds the pen on.

What is still unverified, and what does it cost?

No public source tests whether permission inheritance holds for field-level security and restriction rules specifically, as opposed to record-level sharing. Pricing is a two-invoice structure: Salesforce meters headless consumption through Agentforce Flex Credits, reported at 20 credits per action and roughly $500 per 100,000 credits plus a $125 per user add-on, while Claude inference is contracted separately with Anthropic.

Both gaps point at the same pilot design. Get a sandbox, assign the read-only set to one user, and run Describe against an object with a sensitive field the user should not see. That single test answers the FLS question for your org in an afternoon, and it answers it with evidence rather than a press quote.

On cost, an agentic workflow makes far more API calls than a human clicking through Lightning, and consumption metering stacks on top of per-seat licensing rather than replacing it. Set a credit ceiling and a measurement window before the pilot starts, because the forecastable failure here is a pilot that works perfectly and produces an invoice nobody budgeted.

Most large orgs I work in have a trigger and Flow stack nobody has profiled since it was written. That, rather than the plugin, is what decides whether agent-initiated writes work on day one. The probe above takes ten minutes and tells you which situation you are in.