What happened this week
On 2 August the EU AI Act's Article 50 transparency obligations began to apply in full, the AI Office gained enforcement power over general-purpose model providers, and California's AI Transparency Act became operative. For anyone shipping voice, the operative sentences are that users must be told when they are talking to an AI system rather than a person, and that synthetic audio must carry machine-readable marking.
Two details matter for planning. The Digital Omnibus on AI entered into force on 27 July and defers machine-readable marking to 2 December for systems already on the EU market before 2 August, so anything launched after that date has no grace period. And on 31 July the Commission published the signatory list for the Code of Practice on transparency of AI-generated content: roughly 190 organisations, with Anthropic, Google, Meta, Microsoft, Mistral, OpenAI and Synthesia named among the providers. No pure-play speech vendor appears in the named examples.
The market moved one day early
OpenAI extended SynthID watermarking to audio on 31 July and opened provenance verification through an API. Across the voice-agent platforms, however, no vendor shipped a disclosure prompt, a disclosure toggle, or a disclosure audit log this week. The compliance surface arrived at the model layer, not the layer where most agents are actually assembled.
Models
Grok Voice Think Fast 2.0 is the week's flagship: reasoning in parallel with speech, 0.70s to first audio, and 82.9% on the Artificial Analysis Speech-to-Speech Quality Index. Worth noting for readers who have followed this digest's benchmark coverage: Full-Duplex-Bench now appears as the conversational-dynamics row in a frontier vendor's launch table, where xAI posts 95.1% against 95.7% for GPT-Realtime-2.1. The overall leader is a hair behind on the full-duplex line specifically.
PolyAI's Dialog-RSN-1 proposes a third architecture, fusing turn taking, ASR, and function calling into one audio-native model while leaving TTS outside so enterprises keep their voice. And Qwen Audio Agent is an Apache-2.0 full-duplex runtime that keeps a conversation alive while a coding agent runs tools, which is the clearest reference implementation yet of agent presence during long tool calls.
Turn-taking, from both sides
Last week produced no full-duplex work at all. This week research and platforms converged on the same problem from opposite ends. M3-DuplexBench from NTT is the first multilingual full-duplex benchmark and finds large cross-language gaps. Cocktail-Talker puts respond, listen, and ignore under GRPO. Latent-IM steers conversational moves inside a speech LLM, DuplexGen synthesises scenario-adaptive turn-taking data, and kiloVAD does causal endpointing in 2.1k parameters. Meanwhile Retell shipped mid-call steering and Pipecat 1.7.0 fixed a turn-analyzer bug that fired an inference per transcript fragment. Still absent: any newly named end-to-end full-duplex foundation model.
Also
Prosody-driven jailbreaks show that delivery alone breaks audio LLM safety while the transcript stays clean, which text-only filters cannot see. On the money side, Fish Audio raised $52M, Smallest.ai raised $13M, OVHcloud closed its acquisition of Gladia, and a Munich court ruled against Suno in GEMA's case, the first European decision that training on copyrighted audio abroad can infringe at home.