The center of gravity in AI shifted decisively from chat to autonomy this week, as Perplexity put an always-on agent on a physical Mac mini, DeepSeek shipped a million-token agentic model, and xAI turned Grok into a standing digital coworker. Underneath the product noise the real story is infrastructure: OpenAI unveiled its first custom inference chip while Europe’s champions doubled down on sovereign compute, a signal that the next phase of the race will be won on economics and control, not benchmarks alone.
OpenAI
What happened
OpenAI introduced Jalapeño, its first custom inference chip, promising faster responses, higher throughput and better power efficiency as the design enters deployment planning. The company also pushed its international build-out into Thailand and Brazil and retired the legacy DALL·E GPT inside ChatGPT.
What it means for your agentic build
Custom silicon is OpenAI’s bid to cut the unit cost of inference, the single largest line item once agents run continuously. Expect downstream price and rate-limit relief for high-volume workloads over the coming quarters; plan capacity assuming inference gets cheaper rather than over-architecting around today’s token prices.
Anthropic
What happened
Salesforce and Anthropic announced “Claudeforce,” embedding Claude directly inside the Salesforce CRM, with Salesforce-in-Claude entering open beta in September. Anthropic also made Claude in Chrome generally available on paid plans and widened Claude for Science and Claude for Teachers.
What it means for your agentic build
Claude is moving to where enterprise data already lives: the system of record. If you run Salesforce, native Claude actions shorten the path from model to governed customer data, so evaluate the pilot before building bespoke CRM integrations you may not need.
Google DeepMind
What happened
DeepMind shipped Gemini 3.7 Flash just three weeks after 3.6 Flash, calling it its “most intelligent workhorse model yet for coding and agents,” with gains in software engineering and agentic planning, alongside general availability for 3.5 Flash-Lite.
What it means for your agentic build
The Flash tier is where production agents actually run: cheap, fast, good-enough reasoning. A three-week release cadence means model selection can’t be a one-time decision, so stand up an eval harness that lets you re-benchmark and swap the workhorse model without re-plumbing your stack.
Perplexity
What happened
Perplexity launched Personal Computer, an always-on AI running on a dedicated Mac mini that monitors triggers and executes proactive tasks around the clock, added the ability to start Computer sessions from an email thread, and brought its Comet browser to Enterprise organizations.
What it means for your agentic build
This is the standing-agent pattern: software that acts between prompts rather than only in response to them. Every sensitive action requires explicit approval with a full audit trail and kill switch, and those controls are a useful template for governing your own autonomous deployments.
xAI
What happened
xAI’s Grok 4.6, with a 500k-token context window and configurable reasoning effort, landed on Microsoft Foundry, while Grok Bot graduated beyond beta into SuperGrok and Cursor plans as an always-on AI teammate that works across inboxes, apps and tools.
What it means for your agentic build
Availability on Foundry puts Grok a click away for Azure-committed enterprises, and the Cursor integration shows coding agents consolidating inside the IDE. If your developers already live in Cursor or Azure, you can pilot Grok without opening new vendor paperwork.
DeepSeek
What happened
DeepSeek moved V4-Pro to general availability with an explicit agentic focus, tool use, code execution and multi-step workflows, a one-million-token context window and outputs up to 384k tokens, while raising prices that still undercut Western frontier labs.
What it means for your agentic build
DeepSeek is the price anchor for agentic workloads; even after the increase it remains dramatically cheaper per token than US frontier models. For non-sensitive, high-volume automation it is worth benchmarking as a cost-control option, with the usual governance caveats about where inference runs.
Meta AI
What happened
Meta Superintelligence Labs returned to open weights, releasing Muse Glimmer, a 29.6B-parameter dense multimodal model under Apache 2.0, and open-sourcing Spark 1.2, as Mark Zuckerberg published a manifesto arguing for broad access to the most powerful AI systems.
What it means for your agentic build
A capable Apache-2.0 multimodal model you can self-host changes the build-versus-buy math for regulated and cost-sensitive workloads. If data residency or per-token economics are blocking you, an open-weight base you fully control may now be viable, provided you budget for the MLOps to run it.
Cohere and Aleph Alpha
What happened
Cohere, which acquired Germany’s Aleph Alpha earlier this year, introduced Parse, an enterprise document-intelligence model, and North Mini Code, its first model for developers, deepening a combined European sovereign-AI stack aimed at regulated buyers.
What it means for your agentic build
The merged entity is positioning as the deploy-in-your-environment, data-sovereign alternative to US labs. For European or regulated enterprises, Parse plus on-premise deployment answers the compliance objections that block cloud AI, so put it on the shortlist when sovereignty is a gating requirement.
This Week’s Structural Trends
The standing-agent race is on. Perplexity’s always-on Computer, DeepSeek’s agentic V4-Pro, xAI’s Grok Bot and Gemini’s agent-tuned Flash all point the same direction: software that acts continuously, not only when prompted. The competitive frontier has moved from answer quality to reliable, auditable autonomy.
Distribution is moving to the system of record. Anthropic inside Salesforce, Grok on Microsoft Foundry, Cohere deployed on-premise: the labs are racing to meet enterprise data where it lives instead of asking buyers to pipe it out. Increasingly the integration layer, not the model, is the moat.
Economics and sovereignty are the new battleground. OpenAI’s custom inference chip, Mistral AI’s European Compute Units, DeepSeek’s price anchoring and the Cohere–Aleph Alpha sovereign stack all reflect one shift: as agents run 24/7, cost-per-token and control-of-compute matter as much as raw capability.
Sources
OpenAI News (openai.com/news); Salesforce Newsroom (salesforce.com/news); Anthropic Newsroom (anthropic.com/news); Google DeepMind Blog (deepmind.google/blog); About Meta (about.fb.com/news); xAI Release Notes (releasebot.io/updates/xai); DeepSeek API Docs (api-docs.deepseek.com/news); Mistral AI News (mistral.ai/news); Cohere Blog (cohere.com/blog); Futurum: Cohere Acquires Aleph Alpha (futurumgroup.com); Perplexity Blog (perplexity.ai/hub/blog).

