<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Whisper &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/whisper/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 15:14:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Whisper &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Voice-Controlled AI Agent Lets Quadruped Robot Understand English and Slovak Commands</title>
		<link>https://scienmag.com/voice-controlled-ai-agent-lets-quadruped-robot-understand-english-and-slovak-commands/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:14:45 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced robotics research Slovakia]]></category>
		<category><![CDATA[AI-driven quadruped automation]]></category>
		<category><![CDATA[AI-powered robotic understanding]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[human-robot interaction]]></category>
		<category><![CDATA[intelligent robotic systems]]></category>
		<category><![CDATA[LangGraph]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[multi-modal sensor integration in robots]]></category>
		<category><![CDATA[multilingual command recognition]]></category>
		<category><![CDATA[natural language interface]]></category>
		<category><![CDATA[quadruped robot with natural language processing]]></category>
		<category><![CDATA[quadruped robotics]]></category>
		<category><![CDATA[ReAct agent]]></category>
		<category><![CDATA[robot safety]]></category>
		<category><![CDATA[robotic locomotion control]]></category>
		<category><![CDATA[robotic systems with speech response]]></category>
		<category><![CDATA[ROS]]></category>
		<category><![CDATA[Slovak and English language commands]]></category>
		<category><![CDATA[voice control]]></category>
		<category><![CDATA[voice-controlled AI robot]]></category>
		<category><![CDATA[Whisper]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206235</guid>

					<description><![CDATA[Researchers in Slovakia have built SMaRTAban, a voice-controlled large language model agent that lets a quadruped robot understand spoken English and Slovak commands, interpret its surroundings through vision, and halt safely on operator demand.]]></description>
										<content:encoded><![CDATA[<p>A century ago, the Czech writer Karel Čapek gave the world the word &#8220;robot&#8221; in his play R.U.R., imagining artificial beings that obeyed commands spoken in plain language. Researchers at the Slovak University of Technology in Bratislava, working with the robotics company Panza Robotics, have now brought that vision a step closer to reality. In a study published in the International Journal of Intelligent Robotics and Applications, they present SMaRTAban — the Smart Multi-legged Robotic Transformer Agent Artaban — a system that lets people control a four-legged robot simply by talking to it, in English or in Slovak, while the machine watches, listens, reasons, and answers back.</p>
<p>The platform at the heart of the work is Artaban, a quadruped robot with twelve degrees of freedom, each leg driven by three Maxon motors. It carries four camera modules, two standard RGB and two depth-sensing RGB-D units, plus a time-of-flight camera in its chest, an inertial measurement unit, motor encoders, temperature sensors, and a built-in speaker. A locomotion controller based on Nonlinear Model Predictive Control translates high-level velocity commands into optimized foot trajectories. Onboard computing comes from an Intel NUC 11 with a Core i7 processor and an NVIDIA RTX 2060 GPU running Ubuntu 20.04 and ROS Noetic.</p>
<p>What makes SMaRTAban distinctive is how it stitches together the pieces of a complete voice-command pipeline. Spoken input passes through voice activity detection adapted from the Google Speech Recognition library, then to OpenAI&#8217;s hosted Whisper service, which automatically detects the language and transcribes the audio. The transcript is handed to a ReAct agent built on the LangGraph framework, where GPT-4o reasons step by step and selects from seven custom tools: moving the robot, stopping it, sitting down, standing up, interpreting camera images through GPT-4o vision, synthesizing speech, and waiting. Everything is displayed in a React-based graphical interface so the operator can see every tool call, every response, and every error in real time.</p>
<p>Safety receives unusually careful treatment for a language-driven system. Because large language models are probabilistic by nature, the researchers deliberately avoid relying on the model alone to keep the robot safe. Each tool has a typed schema whose arguments are validated before execution, and invalid parameters are rejected and returned to the model as errors without any physical action. More importantly, nearly every tool runs inside a subprocess guarded by a custom &#8220;interruptible&#8221; decorator. When the operator presses the interrupt button in the GUI, the main process kills that subprocess and independently issues a zero-velocity command to the robot — no LLM consultation required. Shutting down the application triggers the same stop, and a physical emergency-stop button cuts motor power as a final backstop. The researchers are candid that this layered approach still lacks a controller-level watchdog or velocity timeout, so it reduces risk without providing a formal safety guarantee.</p>
<p>The evaluation is unusually thorough for this field. A pilot voice study collected 120 recordings of four phrases — two short and two long, in both English and Slovak — spoken ten times each by three volunteers. English transcription was highly accurate, with word error rates of just 0.04 to 0.08. Slovak proved much harder at the acoustic level: the short phrase &#8220;Choď dopredu&#8221; (go forward) reached a word error rate of 0.97, occasionally being misidentified as a different Slavic language. Yet here is the striking part — even when the transcription looked mangled, meaning often survived. The short Slovak phrase still scored 0.87 on BERTScore, a semantic similarity measure, and the system correctly interpreted the intended command in 87 percent of recordings. The longer Slovak phrase reached a BERTScore of 0.96 and 90 percent understanding despite a word error rate of 0.25.</p>
<p>The core of the paper is a controlled ablation spanning 4,800 accepted tool-calling trials: two models, GPT-4o and GPT-4o-mini, across six agent and prompting configurations, two languages, ten command scenarios, and twenty repetitions each. The scenarios covered translation and turning movements, posture changes, a vision query, and one demanding multi-step command requiring the robot to sit, wait, stand, move forward, turn right, stop, and then describe what it sees. The results tell a clear story about what actually makes language-driven robots reliable.</p>
<p>Simply giving the model access to tools was not enough. Without a task-specific system prompt, GPT-4o completed the full expected action sequence in only 29.75 percent of trials. Adding an explicit system prompt describing the action contract — move, then wait, then stop — lifted that to 52.50 percent. But the decisive jump came from few-shot prompting: three worked examples demonstrating the movement pattern. With iterative ReAct execution, the system prompt, and those demonstrations combined, GPT-4o achieved a tool-call F1 of 0.9977 and 96.5 percent exact-sequence success. GPT-4o-mini followed the same trend, reaching 82 percent exact-sequence success at roughly one-twentieth of the model cost, though it failed to produce a single exact sequence for the complex multi-step command.</p>
<p>That gap between individual correctness and complete sequences is one of the paper&#8217;s most important lessons. On the complex scenario, GPT-4o-mini earned a respectable F1 of 0.8764 because most of its individual tool calls were right — but the ordering was wrong, typically invoking the camera before the required final stop. GPT-4o completed 26 of 40 complex sequences, including all 20 Slovak repetitions. The authors argue, convincingly, that robotics evaluations of language models must report task-level sequence success, not just per-call accuracy, because a command with a missing or misplaced stop action can be physically dangerous even when it scores well on superficial metrics.</p>
<p>Latency measurements complete the practical picture. On the deployed robot&#8217;s own computer, the first tool callback arrived a median of 2.74 seconds after the nominal end of a recorded utterance, rising to 4.10 seconds at the 95th percentile. By contrast, the local ROS stop-service round trip took a median of just 4.28 milliseconds. In other words, cloud transcription and model inference dominate the response path, while local robot communication is nearly instantaneous. This supports a sensible division of labor: the LLM handles interactive, high-level interpretation, while time-critical stabilization and locomotion stay with deterministic local controllers. The study also verified conversation history retention, hazard detection — the system correctly flagged a carpet cutter with an exposed blade as dangerous — graceful error recovery, and reliable mid-execution interruption.</p>
<p>The authors are careful about the limits of their claims. The functional study used scripted commands in dry-run mode, so physical motion quality was not quantified; the voice study involved only three speakers and four phrases; Slovak difficulty at the transcription stage remains an upstream failure mode when meaning is genuinely lost; and dependence on remote APIs introduces network latency and availability concerns. Still, SMaRTAban stands out as the first system in its field to report an evaluated combination of bilingual voice control in English and Slovak, on-demand GPT-4o vision, operator-controlled interruption of running actions, and rigorous quantitative analysis of the whole command path. Future work points toward native multimodal audio input, locally hosted open-source models to remove cloud dependency, the Model Context Protocol for portable tool definitions, integration with SLAM for autonomous navigation, and formal human-robot interaction studies with non-expert users. For now, the image of a quadruped robot that understands &#8220;Choď dopredu&#8221; as readily as &#8220;Go forward&#8221; — and stops the instant you tell it to — marks a meaningful step from Čapek&#8217;s century-old fiction toward an everyday robotic companion.</p>
<p><strong>Subject of Research:</strong> A voice-controlled large language model agent enabling bilingual natural-language control of a quadruped mobile robot with integrated vision and safety mechanisms.</p>
<p><strong>Article Title:</strong> SMaRTAban: a voice-controlled LLM agent for quadruped mobile robotics with integrated vision</p>
<p><strong>Article References:</strong> Zelenay, E., Kocúr, M., Lukáč, M., Duchoň, F., &amp; Marko, R. (2026). SMaRTAban: a voice-controlled LLM agent for quadruped mobile robotics with integrated vision. <em>International Journal of Intelligent Robotics and Applications</em>. <a href="https://doi.org/10.1007/s41315-026-00593-0" rel="noopener noreferrer">https://doi.org/10.1007/s41315-026-00593-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41315-026-00593-0" rel="noopener noreferrer">10.1007/s41315-026-00593-0</a></p>
<p><strong>Keywords:</strong> large language models, quadruped robotics, voice control, Whisper, GPT-4o, ReAct agent, human-robot interaction, natural language interface, LangGraph, robot safety, computer vision, ROS</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206235</post-id>	</item>
		<item>
		<title>New Bilingual Speech Dataset Takes Aim at AI&#8217;s Weakest Spot: Code-Switching</title>
		<link>https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:35:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automatic speech recognition]]></category>
		<category><![CDATA[automatic speech recognition challenges]]></category>
		<category><![CDATA[bilingual speech]]></category>
		<category><![CDATA[bilingual speech recognition]]></category>
		<category><![CDATA[code-switching]]></category>
		<category><![CDATA[code-switching dataset]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[DOTA-ME-CS corpus]]></category>
		<category><![CDATA[improving machine understanding of code-switching]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Mandarin]]></category>
		<category><![CDATA[Mandarin-English code-switching]]></category>
		<category><![CDATA[multilingual natural language processing]]></category>
		<category><![CDATA[multilingual speech processing]]></category>
		<category><![CDATA[open-source language datasets]]></category>
		<category><![CDATA[Paraformer]]></category>
		<category><![CDATA[phonetics]]></category>
		<category><![CDATA[SenseVoice]]></category>
		<category><![CDATA[speech dataset]]></category>
		<category><![CDATA[speech dataset for bilingual speakers]]></category>
		<category><![CDATA[speech recognition for code-switching]]></category>
		<category><![CDATA[transformer-based ASR models]]></category>
		<category><![CDATA[voice conversion]]></category>
		<category><![CDATA[Whisper]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203051</guid>

					<description><![CDATA[A new open dataset of 9300 Mandarin-English code-switched speech recordings, enhanced with AI-generated noise, speed and timbre changes, exposes how badly today's speech recognition models fail at bilingual conversation.]]></description>
										<content:encoded><![CDATA[<p>When bilingual speakers chat with one another, they rarely stay inside a single language. A sentence that begins in Mandarin may slip mid-phrase into English and back again, a behaviour linguists call code-switching. It is one of the most natural things multilingual people do, and one of the most unnatural things for machines to understand. Automatic speech recognition (ASR) systems, even the most powerful transformer-based models now in wide use, tend to stumble exactly at the point where one language hands off to another. A new openly available corpus called DOTA-ME-CS, short for Daily Oriented Text Audio Mandarin-English Code-Switching dataset, has been created to give researchers the fuel they need to close that gap.</p>
<p>The dataset, described in the Journal of Ambient Intelligence and Humanized Computing, contains 18.54 hours of audio spanning 9300 recordings produced by 34 bilingual participants, all of them fluent in both Mandarin and English. Unlike many earlier corpora, every single utterance in the collection involves code-switching. That design choice matters. Older resources such as SEAME, which stretches across roughly 190 hours, and TALCS, which covers 587 hours, contain substantial proportions of monolingual speech, meaning their effective supply of genuinely code-switched material is far smaller than their total length suggests. Some of those datasets, including the ASRU and TALCS corpora, are no longer publicly accessible at all, leaving the field with a shortage of usable, openly available benchmarks.</p>
<p>The construction of DOTA-ME-CS follows an unusual pipeline that blends large language model generation with human recording. The team used GPT-4o with carefully engineered prompts to produce scripted sentences across ten everyday scenarios: education, entertainment, environmental protection, food, health, home, life, pets, travel and work. Each prompt required the model to produce sentences that mimic daily conversational style, contain more English words than Mandarin words, and include at least one Mandarin word, with a dominant language assigned to every sentence. The authors justify this topic-anchored approach with a probabilistic argument: when a specific category is given, the probability of generating a relevant, high-quality sentence is higher than when the model is left to roam across all possible topics, which also reduces hidden cultural bias in the resulting scripts.</p>
<p>Human evaluators checked the generated scripts for grammatical problems and confirmed the presence of genuine switching, and a post-hoc naturalness study asked eight bilingual raters to score 100 randomly sampled stimuli on a five-point Likert scale. The average rating came in at 4.12 with a standard deviation of 0.61, and the median was 4.0, indicating that bilingual listeners generally perceive the generated sentences as natural, though the authors acknowledge that occasional awkwardness is an inherent limitation of LLM-based text generation. Bilingual volunteers, mostly college students with academic backgrounds in China and the United Kingdom, including roughly eighteen participants based at Imperial College London, then recorded the scripts on their own laptops or smartphones as 16-bit WAV files in quiet indoor settings. Recordings that failed basic quality checks were rejected and re-recorded, and accepted files were peak-normalised to 3 dBFS for consistent loudness.</p>
<p>The participants&#8217; linguistic profiles were deliberately diverse. Twelve reported English dominance, eighteen reported Mandarin dominance and four identified as balanced bilinguals, and the mix of speakers from China and the United Kingdom ensures that the corpus spans multiple varieties of English and second-language accents. An average of 2.22 switching points per utterance, combined with a broad part-of-speech distribution across both languages, means the recordings capture switching at varied syntactic positions and grammatical categories rather than concentrating it in a single predictable spot. The dataset also includes longer recordings of roughly 10 to 15 seconds and about 100 words, an intentional response to the weakness current ASR models show on extended speech, where dependencies and critical information can be lost.</p>
<p>What sets the corpus apart most sharply is its AI-driven augmentation. Because human recordings were captured in quiet conditions at a normal pace, the researchers used the Librosa audio library to modify a random subset of clips. Playback speed was adjusted to 0.75x, 0.5x, 1.25x, 1.5x and 2x, with 200 recordings modified at each setting. Five categories of background noise, drawn from highways, war, natural sounds, white noise and playground environments, were added at 200 recordings per type, with the noise pitch scaled to 0.8x to account for the Lombard effect, the well-documented tendency of speakers to raise their vocal intensity in noisy surroundings. In addition, five AI-generated timbres, two male and three female, replaced the original voices in 200 recordings each, simulating both voice conversion deepfake scenarios and privacy-preserving speech processing conditions.</p>
<p>The accompanying data analysis is unusually thorough. Phoneme distributions catalogued with the International Phonetic Alphabet reveal that Mandarin contributes more multisyllabic pronunciations, tones, aspirated consonants and palatalised sounds, while English favours single-syllable vowels; shared features such as the sounds t and a suggest cues that future switching-detection models could exploit. Physical measurements show a frame rate of 46.7 kHz and a sample width of 2.0 bytes, meeting professional audio standards, with an average maximum pitch of 3617.13 Hz and an average per-recording pitch of 535.04 Hz. Formant frequencies computed through Linear Predictive Coding yield averages of 804.02 Hz, 4419.15 Hz and 7549.48 Hz for F1, F2 and F3, pointing to mid-to-low and front vowels, reduced lip rounding and retroflex consonants. The measured speaking rate of 2.05 words per second, with a standard deviation of 0.49, sits squarely within the typical range for conversational read speech, supporting the corpus&#8217;s ecological validity.</p>
<p>To establish baseline performance, the team benchmarked three pre-trained ASR models on the full dataset without fine-tuning: Whisper large-v3, a 1.5-billion-parameter Transformer encoder-decoder trained on 680,000 hours of multilingual audio; SenseVoice Small, a lightweight hybrid CTC-attention model of roughly 200 million parameters that also performs language identification, emotion detection and audio event detection; and Paraformer, a non-autoregressive Transformer with 220 million parameters optimised for Mandarin and built on continuous integrate-and-fire alignment. Paraformer and SenseVoice achieved the best results, with Whisper slightly behind, and a word error rate of 0.177 alongside a character error rate of 0.176. Compared against the publicly available ASCEND corpus, the models produced higher error rates on DOTA-ME-CS, which the authors read not as a defect but as evidence that their dataset poses a more demanding and therefore more useful benchmark.</p>
<p>Case studies expose exactly where today&#8217;s systems fail. In one example embedding the Chinese noun 菠萝包, meaning pineapple bun, inside an English sentence, Paraformer and SenseVoice produced garbled outputs such as bullleball and boobao, while Whisper translated the phrase correctly but destroyed the mixed-language structure by refusing to preserve the code-switch itself, a translation bias that also appeared on noise-free recordings. At double playback speed, Paraformer&#8217;s output strayed entirely from the intended meaning, SenseVoice descended into repetitions such as Miamiami, and Whisper mangled Florida lifestyle into Forensic Lifestyle. Altered timbres degraded English recognition across the board, and all three models performed worst at 2.0x speed. Whisper proved the most resilient to background noise, while Paraformer was the most sensitive to voice changes.</p>
<p>The implications reach beyond engineering. Reliable code-switching recognition would make automated transcription, voice assistants and translation systems far more useful for the hundreds of millions of people who live their linguistic lives between two languages, reducing a bias that currently disadvantages non-monolingual speakers. The authors stress that the study received ethical approval from Imperial College London&#8217;s ethical board, that all participants gave informed consent, and that privacy protections were maintained throughout. They are candid about limitations: funding constrained the scale of the corpus, and scripted reading, while reducing transcription errors and ethical risks, sacrifices the spontaneity of natural conversation. Even so, their deliberate trade-off of quality and coverage over raw scale gives the community something it has lacked, a fully public, exclusively code-switching Mandarin-English benchmark with baseline scores, rigorous acoustic analysis and data, code and recordings all released through a public repository. As speech technology races toward multilingual ubiquity, datasets like this one may determine whether the next generation of voice interfaces can finally follow the way people actually talk.</p>
<p><strong>Subject of Research:</strong> A Mandarin-English code-switching speech dataset for advancing automatic speech recognition research.</p>
<p><strong>Article Title:</strong> DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset</p>
<p><strong>Article References:</strong> Li, Y., Wei, Z., Yu, H., Xue, J., Zhou, H., &amp; Schuller, B. W. (2026). DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05119-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">10.1007/s12652-026-05119-x</a></p>
<p><strong>Keywords:</strong> code-switching, automatic speech recognition, Mandarin, bilingual speech, speech dataset, Whisper, Paraformer, SenseVoice, data augmentation, phonetics, large language models, voice conversion</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203051</post-id>	</item>
	</channel>
</rss>
