Product

Why does yeshcube apply Audio-first in products where voice is the main channel?

Article cover: Why does yeshcube apply Audio-first in products where voice is the main channel?

Audio-first is the design principle through which yeshcube places voice and sound at the center of interaction in products integrating Somia. Screens retain functions involving configuration, review, consent and accessibility when they provide a clearer or more precise representation.

The purpose is to allow an experience to unfold through speaking and listening without requiring continuous visual attention on an interface throughout the session. This priority is particularly consistent with products based on immersive audio, spatial sound and guided conversation. The criterion applies to the design of those products and has no scope as a hub position on digital interaction in general.

An interaction priority

Audio-first determines which channel supports the main experience. It may include spoken language, synthetic speech, music, soundscapes, auditory cues and silence designed as part of the journey.

The graphical interface retains a complementary role. A screen is useful for managing an account, reviewing permissions, comparing options, correcting a transcript or consulting information that needs to remain visible.

This distribution reflects the characteristics of each channel. Visual information can organize several elements in space and keep them available at the same time. Audio unfolds sequentially and requires each fragment to be heard before the next becomes available.

yeshcube applies Audio-first alongside zero-screens and Distraction-Free Apps. The three principles coordinate the sound experience, limited screen use and the design of the applications required to control each product.

Architecture of an Audio-first experience

A conversational voice experience combines several technical and design layers.

  1. Activation. The person starts listening through a voluntary action, physical control, application or wake word, depending on the product.
  2. Audio capture. The microphone records the intervention and communicates clearly when it is active.
  3. Speech recognition. The signal is converted into a transcript or units the system can process.
  4. Interpretation. The architecture identifies intent, available context and the rules that apply to the request.
  5. Orchestration. The system decides whether to respond, request clarification, adapt the experience or apply a stopping rule.
  6. Generation and synthesis. The response is prepared for listening and played through speech and sound.
  7. Repair. The person can correct, repeat, pause or end the interaction.
  8. Supporting interface. The application manages permissions, preferences, privacy and data when the task needs a visual representation.

Quality depends on the entire journey. Natural synthetic speech provides little value when activation is ambiguous, recognition fails or the person has no mechanism for correcting an action.

Benefits in suitable contexts

Rapid language input

Speech can accelerate some text-entry tasks. A controlled study of short messages on mobile phones found input rates close to three times faster through speech recognition than a keyboard in English and Mandarin.

The result applies to laboratory conditions and a specific task. The advantage may decrease with noise, names, numbers, formulae, confidential information or detailed review of the output.

Input speed does not determine total efficiency. Waiting time, clarification and correction are also part of the interaction.

Use when hands or eyes are occupied

Voice allows a person to request an action or receive guidance without continuously operating a screen. This property can be useful in portable devices, acoustic booths, home environments or compatible manual tasks.

Auditory interaction still uses cognitive resources. A hands-free interface can compete with other tasks for attention, particularly when both involve understanding language, making decisions or monitoring the environment.

Access for some people with disabilities

Voice interfaces can make technology easier to access for people who are blind, have low vision or have some motor impairments. They can also reduce dependence on small touch controls and precise gestures.

An auditory experience can create barriers for people who are deaf, have hearing loss, speech difficulties, vocal fatigue or sound sensitivity. Accessible design requires alternative channels suited to each function.

Sound expression

Voice communicates rhythm, emphasis, pauses and turn-taking. These resources help structure a conversation and signal instructions, confirmations or changes within an experience.

Somia adapts interaction to expressed content, available context and authorized preferences. Its current design does not treat tone, pauses or rhythm as a reliable means of determining a person’s emotional state.

Application across the yeshcube ecosystem

Audio-first takes different forms according to the product integrating Somia.

System or productMain role of audioRole of the screenDocumented status
SomiaConversation, content and experience orchestrationAdministration, permissions and reviewERL-3
Somia BloomSpatial sound and guided experiences in acoustic boothsControl from the mobile phoneERL-3
VVAVVEImmersive audio and voice guidance during bounded sessionsSelection and configuration outside the main experienceERL-2
Somia OneHome access through voice and soundLinking and administration through Somia AppERL-0

Somia

Somia provides the shared conversational artificial intelligence core. Its architecture organizes the experience, content, data, analysis and validation.

In Somia Sensory mode, audio plays predefined sound content without conversation. Somia Immersive adds voice interaction and adapts the session to information communicated by the person and the data they have authorized.

Somia Bloom

Somia Bloom brings spatial sound and artificial intelligence into acoustic booths. Bloom Echo uses two speakers in a stereo configuration and Bloom Pulse uses four speakers in a surround configuration.

The booth has no integrated microphones. A mobile phone acts as the control point and voice input when the experience uses conversation. Once started, the session can unfold through the sound system without keeping a display open inside the booth.

VVAVVE

VVAVVE combines immersive audio, gentle light reduction and the Somia architecture in a portable mask. Sessions have a defined beginning and ending and use sound, voice and silence as their main elements.

The visual infrastructure supports selection and administration. During the session, the screen ceases to occupy the center of the interaction.

Somia One

Somia One is a personal home device using voice and sound and linked to Somia App. Its design includes a wake word, a visual indication of microphone status and physical privacy controls.

The application manages the account, language, connectivity and permissions. Everyday use is planned around conversation and audio playback. The product is at ERL-0, with its functional architecture defined and no validated technical prototype.

Four-quadrant diagram titled "Beneficios del paradigma Audio-first": multilingual support, global expansion, basic interaction and use in specific settings

Relationship with zero-screens

Audio-first assigns the main interaction to sound. Zero-screens reduces the need to keep a screen active throughout the experience. Distraction-Free Apps defines how the visual interfaces that remain necessary should operate.

Together they allow an application to act as a control point and move out of the foreground when a session starts. Permissions, privacy and complex information remain available without making the screen the permanent center of the product.

The absence of a screen during an experience does not remove the digital infrastructure. Processing, connectivity, accounts and providers remain part of the system and should be explained clearly.

Limits of the auditory channel

Sequential information

Audio requires information to be received in a temporal order. Comparing several options, reviewing a table, exploring a map or correcting a long document is often more efficient through a persistent visual representation.

Short responses, repetition, playback-speed controls and access to a visual summary can reduce this limitation.

Function discovery

A graphical interface can display available actions through buttons and menus. In a spoken interface, possibilities remain hidden until the system explains them or the person phrases a compatible request.

Research into voice interfaces has found better performance and usability when explicit discovery strategies are provided, including contextual suggestions and help commands.

Recognition variability

Automatic speech recognition performance changes according to language, accent, dialect, age, microphone, noise and speech characteristics.

Research has also identified relevant accuracy differences between groups of speakers. These disparities require systems to be evaluated with the actual populations and conditions of each product rather than relying on a general average.

Privacy

Voice may contain personal information, sensitive data and acoustic characteristics from which inferences can be drawn. Each product should state when capture begins, which device contains the microphone, what information is transmitted and how long it is retained.

Shared spaces create additional exposure risks. A spoken response can reveal information to other people and a voice request may be unsuitable in some environments.

Sensory and communication diversity

An Audio-first interface needs alternatives when a person cannot speak, hear or understand the response under the expected conditions.

Captions, transcripts, tactile controls, vibration and text interfaces may form part of a solution. Their availability should be confirmed for each product rather than attributed automatically to the whole ecosystem.

yeshcube design criteria

yeshcube applies Audio-first through operational criteria.

  1. Bounded purpose. Each experience should define the tasks it handles and the requests outside its scope.
  2. Clear activation. The person should be able to identify when the microphone is active.
  3. Short turns. The main information should appear at the beginning of the response.
  4. Available repair. The interaction should allow correction, repetition, pausing and ending.
  5. Proportionate confirmation. Sensitive actions need an additional check.
  6. Complementary channels. Screens, text and physical controls should remain available when they add precision or accessibility.
  7. Privacy by design. Audio capture and retention should be limited to a stated purpose.
  8. AI transparency. The person should know they are interacting with an artificial system.
  9. Inclusive evaluation. Testing should cover relevant languages, accents, abilities and acoustic conditions.
  10. Product-level validation. Each device and mode needs to establish its own functionality, safety and utility.

Evidence and maturity

Audio-first describes an interaction priority. Applying it to a product does not by itself establish improvements in accessibility, productivity, attention or emotional wellbeing.

Somia and Somia Bloom are at ERL-3, initial transfer. VVAVVE holds ERL-2 based on a controlled pilot with self-reported results and no control group. Somia One remains at ERL-0, with an architecture defined and a prototype still pending.

These levels apply to specific systems and products. Results obtained with a portable mask do not automatically transfer to an acoustic booth or home device.

An Audio-first experience can be evaluated through task completion, recognition errors, repair capacity, understanding of permissions, accessibility, privacy and channel preference in each context.

Interaction that frees the gaze

yeshcube applies Audio-first to the development of products that can accompany an experience without demanding continuous visual attention. Voice and sound allow Somia to extend across booths, portable devices and future home interfaces under a shared architecture.

Screens retain a necessary role in configuration, review and decision-making. The value of the paradigm comes from assigning each task to the most appropriate channel, preserving individual control and evaluating the actual capabilities of each product.

References

Discover our Allies program →

Fill in our contact form →

Other products

Product
Somia™ Bloom
An AI sound system that turns acoustic booths into wellbeing spaces.
Learn more → Somia Bloom
Product
VVAVVE™
A wellbeing mask with emotional AI and immersive audio.
Learn more → VVAVVE
Product
Kiwikidoo™
Screen-free voice adventures so children can create stories and values.
Learn more → Kiwikidoo
Product
Somia™ One
A screen-free home device for everyday emotional companionship.
Learn more → In development
Product
Somia™ Buds
Immersive-sound earbuds for sensory and guided sessions.
Learn more → In development
Product
Somia™ Swish
A voice check-in for older people that alerts the carer if there is no answer.
Learn more → In development
Product
Somia™ Ring
Gesture control for the Somia™ ecosystem, with no phones or apps.
Learn more → In development
Product
Somia™ Niu
Standalone booths that support children through sensory crises.
Learn more → In development
See all products →

Subscribe to our updates

By subscribing, you will receive yeshcube news and content by email. You can unsubscribe at any time. See our privacy policy.

Follow us