otoearth 's Collections

Fullduplex Signals

Weekly signals in speech-to-speech and full-duplex voice AI. Latest: 2026-W35, Aug 17 - Aug 23, 2026. Archive: fullduplex.ai/signals


  • Note 2026-W38 · Delivers fixed, request-independent spoken prompts through the user audio channel while a full-duplex model is answering, comparing fixed-delay interruption with refusal-triggered interruption keyed to the model's streaming text. Across four open models and 720 requests from AdvBench and HarmBench, fixed-delay interruption raises attack success on AdvBench to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, up 33.8 and 39.3 points; the refusal-triggered policy reaches 35.6% and 48.


  • Note 2026-W38 · A controlled ablation of acoustic, prosodic and semantic signals for streaming end-of-turn detection with a lightweight trimodal classifier under identical training. Acoustic plus prosodic features give the best accuracy-latency balance: utterance F1 0.93 with 7.8% false alarms at 400 ms median latency. Adding text increases premature detections without improving performance; prosodic features separate the classes best and text representations overlap. Turn-taking is carried mainly by


  • Note 2026-W38 · Introduces Duplex Cue, an evaluation of the response humans use routinely and full-duplex benchmarks cannot express: continuing to speak while folding in what the listener just contributed, a missing word, a correction or a clarification. It separates listener intent (backchannel, collaboration, interruption) from speaker behaviour (continue, adapt within the turn, yield). A single-model case study uses 300 human-confirmed cues from unscripted English conversation, comparing recorded


  • Note 2026-W38 · A mixture-of-experts, LLM-based ASR system on the Qwen backbone trained on tens of millions of hours, aimed at production gaps rather than benchmarks: regional dialects, dynamic entities and hotwords, long-range context and disfluent spontaneous speech, through a single instruction-following interface. Supports 30 languages and 16 Chinese dialectal varieties across eight dialect regions. Submitted 7 September and revised 9 September. No public weights or repository for this version we


  • Note 2026-W38 · A general-purpose audio generator covering zero-shot TTS, voice design, vocals, sound effects, music, vibe speech and mixtures, built as a discrete autoregressive model over residual vector quantisation tokens rather than the diffusion-transformer approach of most recent general audio models. Its tokenizer represents audio at 12.5 Hz in a shared 16 x 2048 residual code space that jointly quantises semantic and waveform-level features. The backbone predicts the first codebook along tim


  • Note 2026-W38 · Open-sourced 9 September with a technical report (arXiv 2609.08936) the day before. One instruction-driven model covers five task families: speech generation, content editing, enhancement and separation, paralinguistic editing and acoustic editing, trained on about 3.03 billion instruction-audio instances and 1.95 million hours of supervision. A multimodal LLM conditions semantics, a VAE trained on speech, audio and music conditions acoustics, and a hybrid rectified-flow transformer g


  • Note 2026-W37 · A 3B diffusion transformer trained with flow matching on 480k hours of speech, then fine-tuned for cross-lingual dubbing, full-duplex dialogue synthesis and emotional dialogue synthesis. It runs on DAC-VAE latents mapping 48 kHz audio to 25 Hz, over 10x EnCodec's compression, and is alignment-free: alignment learned by cross-attention, no duration predictor. One-shot generation to about a minute, long-form via multi-diffusion. Reported to approach human recordings on short conversatio


  • Note 2026-W37 · Neural finite-state-machine dialogue puts turn-taking control tokens and response text on one causal tape under ordinary next-token prediction, keeping the base LLM's semantics at low fine-tuning cost. Its weakness has been synthetic text training data, since LLMs cannot simulate real acoustic timing. This work learns turn-taking from real human-human spoken dialogue and semantics from human-agent text, with a rule-based transformation that serialises recordings into FSM tapes without


  • Note 2026-W37 · A mechanistic trace of speaking-style information through Whisper-large-v2, Qwen2-Audio-7B-Instruct, Qwen2.5-Omni-7B and Chroma-4B on Expresso, using centered kernel alignment, leave-one-speaker-out probes, open-ended tone prediction and a content-prosody leakage metric. All four strongly encode style in the top third of the audio encoder, and all degrade it before the output. The projector changes geometry without removing information; decoders differ in how much style survives. Mode


  • Note 2026-W37 · Introduces Output Divergence Rate, the share of utterances where speech enhancement changes an LLM's intent classification relative to clean speech, benchmarked over five conditions on 2,974 SLURP clips through Whisper large-v3 and wav2vec2-large cascades. Every condition diverges significantly from zero. MetricGAN+ more than doubles ODR versus unenhanced noisy speech, 0.318 against 0.135, while improving PESQ; unmitigated echo reaches 0.836 through speaker substitution, a failure WER


  • Note 2026-W37 · An LLM-based end-to-end model that produces who-said-what as speech arrives, interleaving fixed-size audio chunks, a small lookahead and previous text so no separate diarization stage is needed. The technical report (arXiv 2609.02812) claims the 7B has the lowest average WER/CER across five evaluation sets and the best or tied-best speaker attribution on 12 of 13 settings. Ten languages and user-supplied hotwords. Weights for both sizes and inference code are MIT. Topped the Hugging F


  • Note 2026-W37 · Six months of real-world usage from more than 500 users of a smartwatch health assistant, yielding 3,030 anonymised utterances that triggered a fallback: noisy audio, transcription errors, ambiguous requests, incomplete utterances and unintended activations. The paper contributes an operational taxonomy, the annotated dataset, and a comparison of classifiers under deployment constraints, finding that lightweight embedding-based classifiers beat larger generative models on most tasks a


  • Note 2026-W37 · 20.0 hours across 58 sessions and seven collaborative tasks (Spot the Difference, Photo Talk, Describe-and-Draw Portrait, Tangram Direction, Consensus Ranking, Hiring Decision, Voice-to-Form), 48 kHz channel-separated FLAC as WebDataset shards, with 11,025 timestamped interface events, per-event visibility, task stimuli and reference answers where the task has one. No transcripts. CC BY 4.0, gated with manual approval. Published 28 August, inside last week's window, and missed there.


  • Note 2026-W36 · A 30-hour hand-labelled corpus of dyadic conversation plus a fixed protocol for scoring end-of-turn and interruption detection, conversation type controlled across six styles, every dialogue triple-annotated at Fleiss's kappa 0.78. Fourteen systems were scored. End-of-turn recall is stable across types; interruption false positives concentrate in backchannel-dense talk. No system is simultaneously fast, high-recall and low on false positives. Disclosure: the training split is oto data


  • Note 2026-W36 · Two omni-modal models converse in native audio with no external speech recognition or synthesis and no API boundary, over the unmodified tasks, tools and success checks of an established text agentic benchmark, so interaction modality is the only variable and the training loop stays local and differentiable. Existing setups either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow, or stay in text and can measure voice agents without improving them. The re


  • Note 2026-W36 · Models how gaze, speech and perceived interpersonal closeness signal floor changes in free four-person dialogue, using the GaMMA corpus and interpretable logistic regression over behaviourally motivated features extracted before each turn-taking event, classifying outcomes as gaps or overlaps. Most turn-taking work this window is dyadic and audio-only; this is the multi-party, multimodal case, and it deliberately trades detector accuracy for features a designer can reason about.


  • Note 2026-W36 · A text-to-speech release covering voice cloning, voice design and voice direction, which reached 215 likes and about 1,800 downloads in its first week and now leads the Hugging Face text-to-speech trending list. The licence is split and worth reading before integration: source code is Apache-2.0, but the model weights, any derivative models, and self-hosted outputs are restricted to research and non-commercial use. BreezeBlue is a new name in this digest and has been added to the org


  • Note 2026-W36 · A streaming Sortformer speaker-diarization and speaker-tagging preview, currently the top trending voice-activity-detection repo on Hugging Face. Access is gated by manual review under the NVIDIA Software and Model Evaluation License: internal test and evaluation only, not production, only on NVIDIA GPUs, no redistribution, no using outputs or artifacts to develop another model, and no disclosure of evaluation or test results without NVIDIA's prior written consent. That last clause ma


  • Note 2026-W35 · 480 persona-grounded scenarios hold the task fixed and vary whether the user's concern is stated in words or carried only in prosody, with objectively checkable outcomes. Giving the model the audio on top of the transcript moves the optimal-solution rate from 14.6% to 15.3%. Forcing it to first write the inferred concern into text takes the same models to 39.6%, against 40.7% for ground-truth state. The prosody is recoverable, and it still does not reach the action unless something ma


  • Note 2026-W35 · Synthesises from uncertain token prefixes instead of waiting for a sentence, using uncertainty-aware buffering and carrying decoder state across segment boundaries. Reported at 15.8ms median time to first token for a single request and 260.8ms at 128 concurrent. Five of its seven authors also wrote X2-Turn, the streaming turn-state model from last week, so one company is now assembling a real-time voice stack part by part without training an end-to-end duplex model at any point.


  • Note 2026-W35 · Instead of sweeping blindly through distortions, this uses cheap structural probes to locate the domain a watermark is embedded in, then applies a single attack matched to that domain, and reports a threshold-free fragility score per scheme. It needs no training and no access to the watermarking model. Anyone planning to satisfy a marking obligation with a neural audio watermark should read it before treating that watermark as the compliance artefact.


  • Note 2026-W35 · Adds real-time predicted backchannels and head nodding to a voice-cloned avatar and measures the effect in a within-subjects study of 35 people. Perceived attentiveness, the sense of talking with the real person, and co-presence all improve significantly. The argument is that a duplex agent feels present because of how it listens, not how well it speaks, which is a case for spending latency budget on the listening side.


  • Note 2026-W35 · Three behavioural probes, covering reference disagreement, masked-number recovery, and orthographic switching, show leading open ASR models reproducing verbatim benchmark reference spans even when the audio contradicts them, masks them, or leaves them ambiguous. The behaviour can be steered with a low-rank direction, which makes it a learned policy rather than an artefact. Anyone choosing a backbone off a WER leaderboard is reading a number that partly measures memorisation.


  • Note 2026-W35 · One backbone with two decoupled continuous paths, an audio encoder for understanding and a RedAE path for generation, covering ASR, audio understanding, zero-shot and instruct TTS, semantic and acoustic speech editing, and temporal grounding over recordings up to an hour. Weights are up under Apache-2.0. The benchmark claims on MMAU, MMSU, Seed-TTS-Eval and InstructTTSEval are self-reported and the linked paper is still a placeholder, so treat the scores as unverified.