Sunday, August 2, 2026

— A weekly publication —

The Agentic Commerce Report

A weekly read of everything that moved in agentic commerce — protocols, payment rails, retailer pilots, regulation. Summarised, sourced, and stitched to what came before.

Anthropic Launches Sonnet 5 and Restores Fable 5 as Multi-Company Jailbreak Framework Proposed

Issue 30June 29 – July 5, 2026Synthesised from 4 sources

Edited by Reviewed against primary sources

Anthropic launched Claude Sonnet 5 on June 30 1, pricing it at $2 per million input tokens and $10 per million output tokens through August 31, after which prices rise to $3 and $15. The model replaces Sonnet 4.6 as the default across Free and Pro plans on claude.ai, Claude Code, and the Claude Platform. On OSWorld-VerifiedA benchmark measuring autonomous desktop and browser control by AI agents, including multi-step task completion and checkout flows. (a benchmark measuring autonomous desktop and browser control), Sonnet 5 posts 81.2%, up from Sonnet 4.6’s 78.5%; on SWE-bench Pro (agentic coding), it reaches 63.2% against Sonnet 4.6’s 58.1% and Opus 4.8’s 69.2%; on GDPval-AAArtificial Analysis's knowledge-work benchmark, scoring models on realistic multi-step professional output tasks. (knowledge work), it scores 1,618 — narrowly above Opus 4.8’s 1,615. Artificial Analysis, which evaluated the model prior to release, found that Sonnet 5 uses approximately 40% more output tokens per task than Sonnet 4.6 and three times the agentic turns on knowledge-work evaluations; at introductory $2/$10 pricing, cost per task runs below Opus 4.8, while at standard $3/$15 pricing it runs approximately 15% higher. Testers at Zapier, Lovable, Kiro, and Pace (an insurance-workflow automation firm) reported that Sonnet 5 completed multi-step agentic tasks — including a Salesforce update combined with enterprise email dispatch in a single autonomous run — where Sonnet 4.6 had not finished.

Sonnet 5 arrives one week after Gemini 3.5 Flash gained computer use and browser control at approximately 30% of GPT-5.5’s cost (2026-w26), and four days after OpenAI previewed GPT-5.6 Sol with restricted access and stronger cyber safeguards. The three releases complete a period in which all three major lab families published a cost-optimized agentic model. Sonnet 5 improves on Sonnet 4.6 across every published benchmark and shows lower hallucination, sycophancy, and prompt-injection susceptibility rates than its predecessor — properties directly relevant to agentic commerce deployments where models interact with untrusted third-party web content and merchant interfaces. Anthropic notes that Fable 5 and Mythos 5 retain higher accuracy on the most demanding tasks; at standard pricing, Sonnet 5 costs less per token than GPT-5.5 and Gemini 3.1 Pro but more than Gemini 3.5 Flash.

US export controls on Fable 5 and Mythos 5, applied June 12 (2026-w24) after Amazon researchers reported a classifier bypass technique, were lifted June 30 2. The US Department of Commerce’s Center for AI Standards and Innovation (CAISICenter for AI Standards and Innovation — a unit of the US Department of Commerce that assessed Anthropic's updated safety classifiers for Fable 5.) reviewed Anthropic’s updated classifiers over two weeks; the revised classifier blocks the reported bypass in over 99% of cases. Commerce Secretary Howard Lutnick signed the reversal; Anthropic committed to pre-release government access, rapid information sharing, and joint research with designated agencies. Fable 5 returned to global access July 1 at up to 50% of weekly usage limits through July 7, transitioning to usage credits thereafter. That same day, Anthropic, Amazon, Microsoft, and Google proposed a four-criterion severity framework for AI jailbreaks 3: capability gain (how far a jailbreak extends an attacker’s access beyond available tools), breadth of that gain, ease of weaponization, and discoverability. Anthropic opened a HackerOne program for Fable 5 vulnerability reporting. The framework formalizes pre-release government testing and interagency vulnerability sharing — no comparable multi-company governance mechanism had previously been proposed at this scope.

Mastercard’s Singapore ‘Next Lap in Payments’ Innovation Circuit opened July 2 4, the fourth edition of the company’s primary Asia Pacific co-creation venue for agentic commerce infrastructure. The Singapore Experience Center — one of seven globally — brings together financial institutions, payment partners, merchants, and policymakers on agentic commerce, trusted digital identity, tokenisation, interoperability, and AI-driven network intelligence. Mastercard confirmed plans for centres in Kuala Lumpur and Tokyo; those markets join Singapore, Malaysia, and Thailand, where live Agent Pay transactions completed on March 4 (DBS, UOB, CIMB, RHB) and April 7 (Krungthai Card) respectively. The three model launches of the past two weeks — Gemini 3.5 Flash with computer use (2026-w26), GPT-5.6 Sol, and Sonnet 5 1 — each extend browser-based agentic execution across price tiers; the Fable 5 restoration 2, the jailbreak severity framework 3, and the Singapore Innovation Circuit 4 mark the governance, security, and infrastructure layers advancing in parallel. No jurisdiction has published regulation specific to agent-initiated commercial transactions.

Events this issue

4 events
Pilots
pilot

Mastercard's Singapore Experience Center showcases agentic commerce tools as company plans centres in Kuala Lumpur and Tokyo

Mastercard's Singapore 'Next Lap in Payments' Innovation Circuit opened July 2; Kuala Lumpur and Tokyo centres planned.

The Innovation Circuit's fourth edition is Mastercard's primary regional demonstration venue for agentic commerce. The Singapore Experience Center, one of seven globally, hosts co-creation sessions with financial institutions, payment partners, merchants, and policymakers across Asia Pacific. The 'Next Lap in Payments' theme covers agentic commerce, trusted digital identity, tokenisation, interoperability, and AI-driven network intelligence — the full stack of Agent Pay infrastructure deployed across the region. Mastercard completed its first live agentic transaction in Singapore with DBS and UOB on March 4, Malaysia with CIMB and RHB the same day, and Thailand with Krungthai Card on April 7. Plans for Kuala Lumpur and Tokyo centres extend the co-creation model to two additional markets, adding the world's third-largest economy to Mastercard's regional agent-payment stakeholder network.

  1. The Asian Banker
Regulation
regulation

Redeploying Fable 5

US export controls on Fable 5 and Mythos 5 lifted June 30; global access resumed July 1 with updated cyber classifiers.

The directive of June 12 (2026-w24) suspended both models for all users worldwide after Amazon researchers reported a technique bypassing Fable 5's safety classifiers. The US Department of Commerce's Center for AI Standards and Innovation (CAISI) reviewed Anthropic's updated classifiers, which block the reported bypass in over 99% of cases. Testing confirmed that Claude Opus 4.8, GPT-5.5, and Kimi K2.7 could identify the same vulnerabilities independently, narrowing the capability-gain assessment. Mythos 5 returned to US organizations under Project Glasswing (2026-w23) from June 26. Fable 5 returns to Pro, Max, Team, and Enterprise plans at up to 50% of weekly usage limits through July 7, then transitions to usage credits. Anthropic committed to pre-release government access, rapid information sharing, and joint research with designated agencies.

  1. Anthropic News
Security
spec

More details on Fable 5's cyber safeguards and our jailbreak framework

Anthropic, Amazon, Microsoft, and Google proposed a four-criterion severity framework for assessing AI jailbreaks.

No consensus definition of AI jailbreak severity exists in the industry. The framework scores a jailbreak on four criteria: capability gain (how far beyond existing tools the jailbreak takes the attacker), breadth of capability gain, ease of weaponization, and discoverability. Anthropic launched a HackerOne program for security researchers to submit Fable 5 jailbreaks for review. The framework builds on a June 2 Executive Order on AI innovation and security (2026-w23) and on CAISI's independent testing of Fable 5's safeguards following the June 12 suspension (2026-w24). A common standard would let model providers triage findings consistently and give governments an agreed threshold for regulatory action. The proposal also formalizes roles for pre-release government testing and interagency vulnerability sharing — the first multi-company agentic security governance framework proposed at this scope.

  1. Anthropic News
Standards
launch

Introducing Claude Sonnet 5

Anthropic launched Claude Sonnet 5 on June 30, pricing it at $2/$10 per million tokens intro rate through August 31.

Claude Sonnet 5 becomes the default model for Free and Pro plans on claude.ai and replaces Sonnet 4.6 across Claude Code and the Claude Platform. Its performance on BrowseComp and OSWorld-Verified matches Opus 4.8 at higher effort levels while costing 33–60% less at standard pricing. Testers at Lovable, Kiro, Pace (an insurance-workflow automation firm), and Zapier reported that Sonnet 5 completed multi-step agentic tasks — including commercial workflows and checkout sequences — where Sonnet 4.6 had not finished. The launch arrives the same week as Fable 5's redeployment (2026-w27) and one week after Gemini 3.5 Flash gained computer use at roughly 30% of GPT-5.5 pricing (2026-w26), adding a third cost-competitive agentic model capable of browser-based commerce tasks within a two-week window.

  1. Anthropic News