OpenAI Previews Ultrafast Mode: GPT-5.6 Sol at 14× Speed
OpenAI made waves on August 13 by previewing Ultrafast mode, a new premium API tier that runs GPT-5.6 Sol at up to 750 output tokens per second — roughly 14 times faster than the standard processing baseline of ~53 tokens per second. The secret sauce? Cerebras wafer-scale hardware, replacing traditional GPU clusters with purpose-built silicon designed for massive parallel inference.
Ultrafast sits above OpenAI's existing Fast mode (announced in late July, offering 2.5× standard speed at 2× the price using GPU infrastructure). The new tier is currently waitlist-only, with early access going to select enterprise customers in coding, commerce, financial research, and customer support. Pricing hasn't been disclosed yet, but OpenAI describes it as a premium tier above Fast mode.
The move signals a broader industry shift: as frontier models converge on capability, speed and latency are becoming the new competitive battleground — especially for real-time agentic workflows where every millisecond counts.
Google Launches Gemini 3.7 Flash With 50% Price Cut
Just hours before OpenAI's Ultrafast announcement, Google launched Gemini 3.7 Flash on August 13 — a fast, efficient multimodal model built for coding, web development, and autonomous agent workflows.
The numbers are impressive: on the DeepSWE v1.1 benchmark for bug-finding and code resolution, Gemini 3.7 Flash jumped from 49.0% to 65.3% compared to its predecessor. Google claims it outperforms comparable models from Anthropic and OpenAI across nine benchmarks.
Perhaps more noteworthy is the pricing: at $0.75 per million input tokens and $3.75 per million output tokens, Google has slashed prices by 50% through the end of 2026. The model is available through the Gemini API, Google AI Studio, Android Studio, and Google's Antigravity agent platform.
The release came just three weeks after Gemini 3.6 Flash, and notably arrives while the much-anticipated Gemini 3.5 Pro remains delayed — a sign that Google is prioritizing rapid iteration on its workhorse models over its flagship.
Google Open-Sources HEIR: AI Inference on Encrypted Data
In a quieter but potentially more consequential announcement, Google released HEIR (Homomorphic Encryption Intermediate Representation) on August 15 — an open-source compiler toolchain that converts pretrained AI models to run inference directly on encrypted inputs.
The technology uses homomorphic encryption, a cryptographic technique that allows computations on ciphertexts without ever decrypting the underlying data. In practical terms, this means a server could run an AI model on your data without ever seeing what that data actually is.
Google's vision is to make HEIR a "one-click solution" for developers who aren't cryptography experts. The release includes demos showing single-threaded CPU latency for applications including recommendation models and hotword detection. The project is part of Google's broader Private Computing Toolkit.
While the technology is still too computationally expensive for many production workloads, the release sparked intense debate on Hacker News and across social media about whether practical encrypted AI inference is finally within reach — a development with enormous implications for healthcare, finance, and any domain where data privacy is paramount.
xAI Ships Grok 4.6: Frontier Parity at Competitive Pricing
Elon Musk's AI company xAI (now branded SpaceXAI) released Grok 4.6 on August 12, just 35 days after Grok 4.5. The new model matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index while keeping the same pricing: $2 per million input tokens and $6 per million output, with cached input at just $0.50.
Key upgrades include a new xhigh reasoning level, improved long-running agent capabilities, and retention of the 500K-token context window. Grok 4.6 scores 1753 ELO and ranks 4th on the AI Intelligence Index, behind Claude Opus 5 but competitive with the rest of the frontier field.
The model shipped simultaneously on Cursor, Grok Build, and the API, with a launch promotion offering 2× included usage inside Grok Build and Cursor for the first week. And xAI isn't slowing down: Grok 4.7, built on a larger 2.1-trillion-parameter architecture, is reportedly weeks away.
EU AI Act Transparency Rules Now Fully in Force
While the tech giants battle over speed and benchmarks, the regulatory landscape shifted on August 2 as the EU AI Act's transparency obligations officially took effect — a milestone that continues to reverberate across the industry two weeks later.
The new rules require AI systems to clearly disclose when users are interacting with AI rather than a human, mandate labeling of deepfakes and AI-generated content with machine-readable markers, and impose transparency requirements around emotion recognition and biometric categorization systems.
Enforcement now sits with national market surveillance authorities and the European AI Office, with penalties of up to €15 million or 3% of global annual turnover. As legal analysts have noted, these obligations are "not delayed, not deferred" — companies operating in the EU must comply now.
The timing is notable: as AI models become faster, cheaper, and more widely deployed, the EU is establishing the first comprehensive enforcement framework for how they're disclosed and used.