Anthropic Removed Forced Tool Use. If Your Callout Pins tool_choice, Fable 5.1 Returns 400.

What changed in the Claude API on 2026-09-01?
Anthropic shipped Claude Fable 5.1 as claude-fable-5-1 on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. Three API changes break existing code: forced tool use is gone, so tool_choice set to "any" or "tool" returns a 400; thinking blocks are model-bound; and editing a prior conversation turn invalidates thinking blocks.
The rest of the release reads like a normal version bump. 1M token context window, 128K max output tokens, always-on adaptive thinking. Token pricing did not move: $10 per million input, $50 per million output, batch at $5/$25. Benchmarks published alongside it include Terminal-Bench-Science 0.1 at 52.6% (Fable 5 scored 24.7%), Terminal-Bench 4.0 at 55.8%, CursorBench 3.2.0 at 73.4%, Humanity's Last Exam at 60.9% without tools and 65.0% with tools, and OSWorld 2.0 at 41.7% strict.
Those numbers are vendor-published and the breaking-change list reached me through a secondary write-up plus the release log rather than the API reference. Confirm the tool_choice behavior against Anthropic's own docs before you put it in a runbook. The audit below is worth running either way, because it costs one command.
Why does removing tool_choice break integrations that work today?
tool_choice: "any" was the cheap guarantee that a model returns a tool call instead of prose. Code built on it skips output validation, reads the tool input straight into a variable, and writes to a record. Without the flag, the same model can answer in plain text, and the parse throws at runtime in production.
This parameter shows up specifically in the integrations somebody thought carefully about. Nobody sets tool_choice on a chat feature. You set it when a field on a record depends on the model returning a specific JSON shape, and prose in that slot means a null pointer or a corrupted value. The flag was doing the job of a schema validator, which is exactly why its removal hurts: the teams that used it are the teams that decided they did not need a validator.
The failure timing is the bad part. A removed request parameter is not a compile error and not a deploy warning. Your Apex compiles, your test classes mock the callout and pass, your metadata deploys clean. The 400 arrives the first time a real user triggers the real path against the new model.
Where does tool_choice hide in a Salesforce org?
In custom Apex callouts to the Claude API or Bedrock, Flow HTTP callout request bodies, External Services registrations, Node or Python middleware sitting between the org and the provider, and MCP server configs. Agentforce's own model invocations are Salesforce-managed, so those are not your exposure. Your surface is every place your code composes the request body.
Run this against the repo before anything else:
grep -rn "tool_choice" --include=*.cls --include=*.flow-meta.xml --include=*.js --include=*.py --include=*.json .Then check the places a repo grep cannot see. Request bodies typed directly into a Flow HTTP Callout action live in the Flow XML, so the flow-meta.xml include catches most of them, but an External Service schema registered from a URL does not live in your repo at all. Neither does a config file on whatever host runs your middleware. In a large org the same JSON body is usually pasted into three or four integrations by three or four different people, and the grep finds the ones under version control, not the ones under someone's laptop.
What replaces the guarantee once the flag is gone?
Validate the response shape in your own code. Check that a tool_use block came back, assert the required keys are present and non-null, retry once with the rejection reason appended to the prompt, then fail deterministically. Two attempts is enough. A model that ignores a direct instruction twice will not comply on the third.
public with sharing class StructuredCall {
public class ShapeException extends Exception {}
// ponytail: fixed 2 attempts, no backoff. Add backoff only if attempt 2
// starts failing measurably in production.
public static Map<String, Object> invoke(String prompt, Set<String> requiredKeys) {
String lastError;
for (Integer attempt = 0; attempt < 2; attempt++) {
String body = (attempt == 0)
? prompt
: prompt + '\n\nYour previous reply was rejected: ' + lastError
+ '\nReply with the tool call only, no prose.';
HttpResponse res = ModelGateway.send(body); // your existing callout
try {
return requireKeys(ToolCallParser.firstToolInput(res), requiredKeys);
} catch (ShapeException e) {
lastError = e.getMessage();
}
}
throw new ShapeException('No valid tool call after 2 attempts: ' + lastError);
}
private static Map<String, Object> requireKeys(
Map<String, Object> input,
Set<String> required
) {
if (input == null) {
throw new ShapeException('Model replied with text, no tool_use block');
}
for (String k : required) {
if (!input.containsKey(k) || input.get(k) == null) {
throw new ShapeException('Missing required key: ' + k);
}
}
return input;
}
}Feeding the rejection reason back into attempt two is the part that earns its keep. A bare retry sends the identical prompt and usually gets an identical answer. Telling the model which key was missing changes the input, which is the only thing that changes the output.
Two operational notes. Each attempt is a separate HTTP callout, so a synchronous Apex transaction gets 100 callouts and 120 seconds of cumulative callout time to work with, and a two-attempt call now costs up to two of them. And log which attempt succeeded. If attempt two starts carrying a meaningful share of your traffic, your prompt is the problem, not the model.
What does the 75% cache-read cut change in a cost model?
Cache reads dropped from $1.00 to $0.25 per million tokens in the same release. Input stays at $10/M and output at $50/M. Any pattern that re-sends a large static prefix on every turn, org schema, Knowledge articles, prompt templates, gets 75% cheaper with zero prompt rewriting. The saving scales with turns per session, not with users.
Concrete shape. Say a 40,000-token static prefix, 12 turns per session, 11 of those turns reading from cache. That is 440,000 cache-read tokens per session. At 2,000 sessions a month you are reading 880 million tokens: $880 at the old rate, $220 at the new one. Same prompt, same architecture, $660 a month back.
The number that decides whether any of this reaches an Agentforce workload is the Bedrock line, because Claude runs in the Salesforce VPC through Bedrock. Fable 5.1 is listed as available there on day one, which is the unusual part. Model availability inside a managed platform normally lags the API by weeks.
When should you stay on the model you are running?
Stay put if you have a working pinned integration, no cost pressure, and no eval harness to prove the replacement behaves the same. A version bump that removes a request parameter is a migration, not an upgrade, and doing it without a way to measure the before and after means finding out from a user.
The move that pays regardless of whether you switch is decoupling the validation from the vendor flag. tool_choice disappeared on somebody else's schedule. The next parameter will too. Code that checks its own inputs does not care which model answered, and that check is thirty lines you write once.
Most of the orgs I work in have this exact shape: an LLM integration written by a consultant who left, a model string and a request body pasted across four places, and nobody currently on the team who knows all four. The grep takes ten seconds and tells you whether you have a countdown running.
