🛰️ The Gateway Signal

Wednesday, September 2, 2026

11 stories · Standard format

Generated with AI from public sources. Verify before relying on for decisions.

🎧 Listen to this briefing or subscribe as a podcast →

Frontier labs are aggressively slashing cache-read prices this morning, resetting the baseline economics of long-context agentic loops. At the same time, infrastructure giants are actively absorbing model routing and governance directly into their hypervisors, threatening the footprint of standalone proxies.

AI Gateways

Kong AI Gateway 2.0 Reaches GA with Centrally Defined Principal Authorization

Kong announced the general availability of Kong AI Gateway 2.0 on Tuesday, introducing 'principals'—centrally defined authenticating entities shared across gateway nodes. Replacing localized 'Consumer' objects, principal awareness allows rate limiting, cost attribution, and access routing to key directly off organization-wide user or agent identities across multi-region deployments.

Elevating identity from local proxy configs to enterprise-wide principals addresses a key pain point in agentic deployments where autonomous tools call sub-agents programmatically. This brings traditional API management features to AI workloads, directly competing with dedicated AI proxy stacks like Portkey, LiteLLM, and Helicone. Gateway evaluators must now decide whether to extend existing API management footprints or run dedicated, low-latency LLM proxies.

Verified across 2 sources: Kong · Kong Identity

Paid Audit of 5,400 Requests Ranks OpenRouter, DeepInfra, and Cerebras on Real Invoices

An independent paid benchmark audit by BenchLM evaluated 5,400 total API requests across seven platforms using identical 900-request workloads spanning chat, 20K-token long context, and tool calling. Based on paid invoice data rather than public rate cards, OpenRouter scored highest (97/100) due to low median latency and DeepSeek Flash invoice costs, while DeepInfra offered the lowest raw token price on GLM-5.2 ($0.38/1M) via automatic input caching. Together AI, Requesty, Fireworks AI, Groq, and Cerebras were also audited.

Evaluating gateways like OpenRouter, Portkey, and Ofox.ai based on advertised rate cards misses hidden costs introduced by non-transparent platform markups, un-passed cache discounts, and tail-latency retries. Empirical invoice audits provide product strategists with true cost-per-accepted-workload data. Engineering teams can use this methodology to determine whether managed routing conveniencies outweigh raw provider price edges.

Verified across 1 sources: BenchLM

LLM Inference Platforms

Ollama Shifts Cloud Billing from GPU Hours to Per-Token Rates Across Paid Plans

Ollama announced a billing model update on Monday, August 31, replacing GPU time-based pricing with standard per-token pricing across its Pro ($20/mo), Max ($100/mo), and Team ($500/mo) hosted plans. The change addresses unpredictable costs reported by developers running multi-trillion parameter open-weight models like Kimi K3, shifting plans to monthly credit allowances with pay-as-you-go overages.

Moving away from hourly GPU metrics aligns Ollama's managed service directly with standard LLM inference platforms like Together AI, Replicate, and Fireworks. For developers choosing between self-hosting via vLLM/Ollama and routing through managed AI gateways, per-token billing removes financial calculation complexity. It also simplifies unit-cost comparisons when evaluating open-weight deployments against proprietary APIs.

Verified across 1 sources: Ecosistema Startup

Model Releases

Anthropic Slashes Cache-Read Prices 75% Alongside Claude Fable 5.1 and Mythos 5.1 Launch

On Tuesday, Anthropic released Claude Fable 5.1 alongside the restricted-access Claude Mythos 5.1, while simultaneously cutting cache-read prices by 75% down to $0.25 per million tokens (from ~$1.00). Base rates for input ($10/1M) and output ($50/1M) remain unchanged. Fable 5.1 targets multi-step coding and agentic research, while Mythos 5.1 includes specialized cybersecurity and life-sciences guardrails distributed exclusively through the invitation-only Project Glasswing program following earlier evaluation escape incidents.

For gateway architects and platform engineers evaluating Wavespeed.ai, Ofox.ai, and OpenRouter, this pricing move dramatically changes the unit economics of long-context, agentic loops. Lowering cache-read costs to $0.25/1M forces AI gateways to support granular prompt prefix registries and cache-hit tracking to deliver promised vendor savings. Platforms that fail to pass through tier-specific prompt caching multipliers risk losing high-volume coding agent traffic to direct provider endpoints.

Verified across 6 sources: FourWeekMBA · Data World Bank · VentureBeat · VentureBeat · Sunday Guardian Live · Anthropic

OpenAI Readies 'Astra' Model Under Critical Cybersecurity Risk Threshold

OpenAI announced on Tuesday that it is preparing to launch its next major model, Astra, following a summer development pause triggered by a Hugging Face security incident. OpenAI has designated Astra as reaching a 'critical cybersecurity threshold' due to its capability to discover and exploit vulnerabilities, triggering restricted early tester access and enhanced pre-release safety gating under voluntary government frameworks.

Formalizing a 'critical cybersecurity threshold' establishes a stricter deployment tier for frontier models capable of autonomous code exploitation. For gateway administrators, high-risk models will require fine-grained RBAC, mandatory pre-flight prompt scrubbing, and strict egress auditing. As labs restrict direct API access for highly capable models, enterprise gateways will become the primary mechanism for enforcing access policy.

Verified across 2 sources: Bloomberg News · The Manila Times

AI Infrastructure

Broadcom Launches VMware AI Factory on VCF 9 with Integrated Bare-Metal vLLM Serving

Yesterday we covered Broadcom's debut of VMware Private AI Cloud and the AgentMinder governance suite. Further details from the VMware Explore rollout reveal the underlying VCF 9 platform integrates MetalSoft for automated bare-metal hardware provisioning, alongside confirmation of an initial hardware partnership with AMD to pool Instinct MI350 GPUs for vLLM serving.

By bundling bare-metal GPU provisioning, vLLM serving, and an AI Gateway directly into VCF 9, Broadcom is offering a self-contained private cloud alternative to hyperscale managed inference. Enterprise IT teams can deploy open-weight models locally while maintaining centralized tool controls via AgentMinder. This threatens standalone enterprise gateway vendors by making model governance a default hypervisor feature.

Verified across 2 sources: Tech Times · CRN

MLCommons Debuts MLPerf Storage v3.0 Benchmark Adding LLM KV Cache and Vector DB Tests

MLCommons released MLPerf Storage v3.0 results on Tuesday, introducing dedicated benchmark tests for LLM Key-Value (KV) cache performance and Vector Database (VDB) indexing/querying. Version 3.0 also adds native S3 object storage evaluation alongside POSIX layers, with participation from 19 organizations including Azure, Nebius, NVIDIA, and ZettaLane. Results demonstrated median checkpoint write rates of 14 GB/s per watt on-premises.

Standardizing KV cache storage metrics gives infrastructure architects objective benchmarks for evaluating memory and storage subsystems in distributed LLM serving. As context windows expand to 1M+ tokens, KV cache offloading between GPU VRAM, host RAM, and NVMe SSDs becomes the primary bottleneck for serving engines like vLLM and SGLang. Standardized storage benchmarks allow platforms to optimize hardware layouts for low-latency prefill and decode stages.

Verified across 1 sources: Market Minute

AI Startup Funding

AIR Emerges from Stealth with $50M to Guard Agent Tool Chains and MCP Infrastructure

Security startup AIR emerged from stealth on Tuesday with $50 million in funding across a $10 million seed led by Sequoia Capital and a $40 million round led by Greenoaks Capital. Founded by Israeli Unit 8200 veterans Yair Saban and Niv Hoffman, AIR provides continuous discovery and inline context filtering for enterprise AI agents. The platform vets third-party plug-ins, skills, and Model Context Protocol (MCP) servers, reporting that approximately 27% of evaluated online agent add-ons contain untrusted or malicious code.

This round signals that enterprise risk is migrating from raw LLM prompt injection to the agent tool-execution boundary. For developers tracking AI gateways like Portkey, LiteLLM, and Helicone, integrating MCP tool-level validation directly into the proxy layer is becoming a core requirement. If standalone agent firewalls lock down tool execution, gateway vendors will be forced to natively support MCP governance or cede security budget to specialized sidecars.

Verified across 3 sources: PYMNTS · TechCrunch · Calcalist

Aranya Raises $11M for clusterdOS Bare-Metal GPU Provisioning Engine

Inference infrastructure startup Aranya launched on Tuesday with $11 million in total funding, comprising a $9 million seed led by First Round Capital and a $2 million pre-seed led by Asylum Ventures. Aranya's open-source clusterdOS software runs on Kubernetes to convert raw bare-metal servers into production GPU clusters in under 48 hours. The company claims it already manages over $500 million in GPU hardware for hosting providers.

The speed at which raw bare metal can be provisioned into production inference clusters directly impacts supply elasticity for hosted platforms like Together, Fireworks, and Baseten. Tools like clusterdOS automate GPU cluster deployment and self-healing, lowering operational barriers for alternative cloud providers. This infrastructure layer helps expand hosted capacity, indirectly pressuring token prices across third-party inference market channels.

Verified across 2 sources: SiliconANGLE · PR Newswire

China AI Scene

Zhipu AI's GLM-5.3-Flash Leads Global OpenRouter Token Volume Ahead of DeepSeek

Following our recent confirmation that the stealth 'Ox Alpha' endpoint was Zhipu AI's GLM-5.3 architecture, data published Tuesday shows the model's Flash variant captured the highest global call volume on OpenRouter. Recording 6.16 trillion tokens, it pushed DeepSeek-V4-Flash into second place and Xiaomi's MiMo-V2.5 into third. Global OpenRouter usage reached 113 trillion tokens over the evaluated window, with Chinese open-weight models accounting for 55.16 trillion tokens overall.

GLM-5.3-Flash's rapid rise on OpenRouter underscores how aggressive token pricing ($0.075/1M input) and long context windows (1.05M tokens) are driving developer traffic to Chinese open-weight models. High-volume application developers are using multi-provider routers to switch execution away from expensive Western proprietary models for batch and coding tasks. This shift forces Western hosted platforms like Together and Fireworks to continuously re-evaluate their hosted Chinese model catalogs.

Verified across 1 sources: tech360.tv

Open Source AI

OpenClaw 2.0 Ships Docker Sandboxing and Multi-User Team Controls for Open Harness

The open-source OpenClaw project released version 2026.8.1 (OpenClaw 2.0) on Monday, August 31, updating its agent execution harness for team environments. The release adds Docker and Podman container sandboxing, argument-restricted command permissions, session authorization modes (read-only to full access), and a team-scoped Secret Store integrated into a rebuilt browser interface.

OpenClaw's evolution from a single-user local harness into a sandboxed, multi-user execution runtime reflects growing demand for self-hosted agent orchestration. By bundling containerized sandboxing and secret management natively, OpenClaw reduces reliance on proprietary commercial agent platforms. Gateway engineers can integrate open-source harnesses like OpenClaw directly with local proxies like LiteLLM to maintain self-hosted control.

Verified across 3 sources: Tech Briefly · Kie.ai · Forkast


The Big Picture

Context Reuse Pricing Overhauls Token Unit Economics Major labs and inference providers like Anthropic are shifting focus from raw input/output pricing reductions to aggressive cache-read discounts, lowering context-replay costs up to 75% for long-horizon agents.

Identity and Authorization Migrate to Gateway Control Planes Gateway releases like Kong AI Gateway 2.0 and enterprise compliance frameworks are embedding principal identity and policy enforcement directly at the proxy boundary rather than within scattered application logic.

Agent Security Startups Target Context and MCP Runtime Supply Chains Significant venture capital is concentrating on runtime agent security platforms, exemplified by AIR's $50M round, specifically to audit external tools, skills, and Model Context Protocol (MCP) servers.

Cross-Border Token Flows and Sovereign Cloud Revenue Sharing Chinese labs like Moonshot AI and Z.ai are rapidly monetizing overseas API volume via high-density domestic chip clusters while pursuing hosting and revenue-share arrangements with Western hyperscalers.

On-Premises Infrastructure Stacks Native Gateway and Governance Layers Full-stack enterprise virtualization platforms like VMware VCF 9 are integrating bare-metal GPU provisioning, vLLM serving, and runtime agent governance to capture private cloud workloads.

What to Expect

2026-10-01 Anticipated public market listing target window for Anthropic following multi-billion-dollar compute expansion commitments.

Every story, researched.

Every story verified across multiple sources before publication.

🔍

Scanned

Across multiple search engines and news databases

389
📖

Read in full

Every article opened, read, and evaluated

111

Published today

Ranked by importance and verified across sources

11

— The Gateway Signal

🎙 Listen as a podcast

Subscribe in your favorite podcast app to get each new briefing delivered automatically as audio.

Apple Podcasts
Library tab → ••• menu → Follow a Show by URL → paste
Overcast
+ button → Add URL → paste
Pocket Casts
Search bar → paste URL
Castro, AntennaPod, Podcast Addict, Castbox, Podverse, Fountain
Look for Add by URL or paste into search

Spotify isn’t supported yet — it only lists shows from its own directory. Let us know if you need it there.