, a 窪蹋勛圖厙 doctoral student, recently toured a museum in Mexico. Chen doesnt speak Spanish, so he ran a translation app on his phone and pointed the microphone at the tour guide. But even in a museums relative quiet, the surrounding noise was too much. The resulting text was useless.
Various technologies have emerged lately promising fluent translation, but none of these solved Chens problem of public spaces. , for instance, function only with an isolated speaker; they after the speaker finishes.
Now, Chen and a team of 窪蹋勛圖厙 researchers have designed at once, while preserving the direction and qualities of peoples voices. The team built the system, called Spatial Speech Translation, with off-the-shelf noise-cancelling headphones fitted with microphones. The teams algorithms separate out the different speakers in a space and follow them as they move, translate their speech and play it back with a 2-4 second delay.
The Apr. 30 at the ACM CHI Conference on Human Factors in Computing Systems in Yokohama, Japan. The code for the proof-of-concept device is available for others to build on. Other translation tech is built on the assumption that only one person is speaking, said senior author , a 窪蹋勛圖厙 professor in the Paul G. Allen School of Computer Science & Engineering. But in the real world, you cant have just one robotic voice talking for multiple people in a room. For the first time, weve preserved the sound of each persons voice and the direction its coming from.
Related:
- Story in
- For more information, visit
The system makes three innovations. First, when turned on, it immediately detects how many speakers are in an indoor or outdoor space.
Our algorithms work a little like radar, said lead author Chen, a 窪蹋勛圖厙 doctoral student in the Allen School. So its scanning the space in 360 degrees and constantly determining and updating whether theres one person or six or seven.
The system then translates the speech and maintains the expressive qualities and volume of each speakers voice while running on a device, such mobile devices with an Apple M2 chip like laptops and Apple Vision Pro. (The team avoided using cloud computing because of the privacy concerns with voice cloning.) Finally, when speakers move their heads, the system continues to track the direction and qualities of their voices as they change.
The system functioned when tested in 10 indoor and outdoor settings. And in a 29-participant test, the users preferred the system over models that didnt track speakers through space.
In a separate user test, most participants preferred a delay of 3-4 seconds, since the system made more errors when translating with a delay of 1-2 seconds. The team is working to reduce the speed of translation in future iterations. The system currently only works on commonplace speech, not specialized language such as technical jargon. For this paper, the team worked with Spanish, German and French but previous work on translation models has shown they can be trained to translate around 100 languages.
This is a step toward breaking down the language barriers between cultures, Chen said. So if Im walking down the street in Mexico, even though I dont speak Spanish, I can translate all the peoples voices and know who said what.
, a research intern at HydroX AI and a 窪蹋勛圖厙 undergraduate in the Allen School while completing this research, and , a 窪蹋勛圖厙 doctoral student in the Allen School, are also co-authors on this paper. This research was funded by a Moore Inventor Fellow award and a .
For more information, contact the researchers at babelfish@cs.washington.edu.泭