— A weekly publication —
A weekly read of everything that moved in agentic commerce — protocols, payment rails, retailer pilots, regulation. Summarised, sourced, and stitched to what came before.
Anthropic launched Claude Opus 5 on July 24 1, priced at $5 per million input tokens and $25 per million output tokens — the same price as its predecessor, Opus 4.8. On OSWorld 2.0Anthropic's version of the OSWorld computer-use benchmark, measuring autonomous agent performance on checkout and multi-step desktop and browser tasks. (the autonomous computer-use benchmark covering checkout and multi-step agent task completion), Opus 5 surpasses Fable 5’s best result at one-third the cost. On Zapier AutomationBenchA benchmark measuring whether AI models complete full business tasks from a stated goal, covering CRM updates, email dispatch, and account management workflows. (a benchmark testing end-to-end business workflow completion), Opus 5’s pass rate runs 1.5 times higher than the next-best model at equivalent cost per task 1. Google released Gemini 3.6 Flash three days earlier 2 at $1.50 input and $7.50 output per million tokens — the new default Flash model in the Gemini family. Gemini 3.6 Flash scores 83.0 on OSWorld-VerifiedGoogle's standardized benchmark measuring AI agents' ability to control desktop and browser interfaces, including multi-step checkout flow completion. (the autonomous desktop and browser control benchmark), up 4.6 points from 3.5 Flash, using 17% fewer output tokens per agentic session 2.
Opus 5 arrives four weeks after Anthropic launched Claude Sonnet 5 (2026-w27), which replaced Sonnet 4.6 as the default across Free and Pro plans and scored 81.2% on OSWorld-Verified. That release followed Gemini 3.5 Flash gaining native computer use (2026-w26), which added the capability at 30% of GPT-5.5’s cost. 2026-w30 marks the first week in which Anthropic and Google both released new model-family entries on overlapping dates: Gemini 3.6 Flash on July 21, Opus 5 on July 24. Both launches add computer-use results from different evaluation suites to the same comparison table for the first time. Opus 5’s Frontier-Bench v0.1An evaluation measuring AI model performance on real-world production-grade software engineering tasks using a standardized agent harness. result (an evaluation measuring performance on production-grade software engineering tasks) more than doubles Opus 4.8’s score at lower cost per task 1. Gemini 3.6 Flash’s DeepSWE score (an agentic coding evaluation) rises from 37% to 49% over its predecessor 2.
Michaels launched Ask Mike on July 21 4, an AI shopping assistant on Google Cloud’s Gemini Enterprise for Customer ExperienceGoogle Cloud's managed Gemini deployment for retail applications, with enterprise access controls and SLA guarantees for production operators. (a managed Gemini service for enterprise retailers). Ask Mike logged 75,000 customer conversations since going live on michaels.com in May 2026; more than 60% of those sessions focused on product discovery. Conversion rates doubled versus traditional keyword search, Michaels reported. The deployment adds craft retail to the set of verticals with named AI shopping-agent pilots in 2026. Prior additions include grocery (Amazon’s Alexa for Shopping, 2026-w06), quick-service restaurants (Papa Johns and Gemini, 2026-w26), and European grocery (Carrefour and ChatGPT, 2026-w13). Plans include expanding Ask Mike with AI-generated product overviews and contextual prompts on individual product pages.
Google announced on July 24 3 that it is signing the EU AI Act Code of Practice on Transparency of AI-Generated Content, committing to mark AI-generated output as machine-detectable. Article 50 of the EU AI Act, which governs the code, takes effect August 2, 2026. Violations carry fines of up to €15 million or 3% of worldwide annual turnover. Apple, ElevenLabs, Kakao, NVIDIA, and OpenAI joined the commitment through SynthIDGoogle's content-watermarking technology that embeds machine-readable provenance signals in AI-generated text, images, audio, and video. watermarking (a technology that embeds machine-readable signals in AI-generated text, images, audio, and video). Adherence is voluntary, but non-signatories must demonstrate compliance through other means and face greater regulatory scrutiny. General-purpose AI model obligations under the EU AI Act also take effect August 2, extending the Act’s scope to foundation models. No jurisdiction has enacted regulation specific to agent-initiated commercial transactions as of this issue.
Agentic commerce is the practice of AI agents initiating, negotiating, and completing commercial transactions autonomously — browsing inventory, comparing prices, applying promotions, and executing checkout without per-step human approval. The agent acts on a consumer's standing intent, constrained by a pre-authorised budget and a defined preference set. McKinsey estimates agentic AI will automate workflows touching $3–5 trillion in commerce annually by 2030. This publication tracks seven lanes where agentic commerce is advancing in real deployments: payment rails (how agents settle transactions), AEO and discovery (how agents find and rank products), standards and protocols (the open specs agents use to interoperate), identity and trust (how agents prove authority to act), security and risk (the new attack surfaces agents create), regulation (how governments and central banks are responding), and retailer pilots (live deployments and their reported results). Every event cited here links to a primary source.
Read the full guide: agentic commerce explained →