Calls tend to slow down at the same point, when somebody starts reading information out loud, whether that is appointment windows, flight times, or a confirmation code. The customer asks for it again, then asks a second time, and ends up writing it down on whatever is nearby. Nobody does that with a text message, because the text is still on the screen when they look back.

Speech is gone once it is spoken. Choices, figures, codes and anything the customer has to act on later belong in a message, and reading them out loud is a design mistake.

A better voice does not make a list easier to follow

Read a customer two flight options in full, a 7:15 departure from Miami landing at LaGuardia at 9:20 with a 4:12 return that gets back at 9 after a layover in Charlotte, against a 9:15 departure with its own return times, and they have lost the first one before the second is finished. Two options are no easier than six, because the customer still has to hold all of the first one in their head while the second one arrives. The usual advice is to read fewer options at a time, and researchers who tested it found that it does not help.

Texted to the customer during the call, those same two itineraries cost almost nothing to read, and neither would six. The customer takes them in at a glance, at their own pace, and still has them after the call. A better model or a warmer voice does not change any of that.

Saying it again often produces the same error

In a classic study of speech-interface errors, researchers at Carnegie Mellon found that when spoken input fails, repeating yourself tends to fail the same way (the same misrecognition recurs), while switching to another modality corrects errors far better than trying the ear again. Users figure this out on their own: they start with speech, hit the wall, and learn to switch. It takes them a while, though, so a well designed conversation offers to text the customer at the point where the information is better read than heard, rather than waiting to be asked.

Every customer who has said “P as in Paul” has hit that wall. Spelling a name aloud, reading back a confirmation code, describing what a damaged bumper looks like: each one makes the customer work at something a photo or a text message would carry effortlessly.

Speaking is still the fastest way to describe a problem.

None of this makes speech a lesser channel. Most people can describe a problem out loud faster and more completely than they can type it, and speaking works when their hands and eyes are busy, which is often why they called instead of opening an app.

So speech should carry the conversation: what the customer wants, what went wrong, and what they decide. Everything else goes to the screen, which is what a good human agent is doing when they say “I’ll send that over to you.”

How Quiq splits the work between the call and the thread

Quiq’s Voice AI sends and receives messages during the call, because the call and the message thread are the same conversation. The AI agent sends the two itineraries as a carousel and keeps talking while the customer looks at them. A photo the customer texts back arrives in that same conversation, and the AI agent can act on it before the call ends.