Abstract:
In the past few days, everyone must have been annoyed by the new function of WeChat grayscale testing. First, in voice input mode, the border of "press and hold to talk" suddenly becomes thick and black, for fear that you can't see it. Then in the text input mode, the input box actually reminds you:
You can press and hold to convert text.


What does this mean?
In the past, "Press and hold to speak" required you to switch to voice mode before you could use it. What you let go was the original voice, and what was heard on the other side was your damn charming bubble voice (unless you actively chose to convert it to text).

But now, you just need to hold down the input box to speak, and the voice will directly turn into text, and it will be sent out when you let go.

At first glance, it seems very convenient. You can send it directly after saying it. However, the first reaction of many netizens is to complain, and it is too easy to accidentally trigger the social death.

But the bad reviewer is not complaining about this today.
Have you noticed that there are too many places for voice input in this app?
You must know that some time ago, WeChat just put a microphone button (+1) on the right side of the input box, and you can convert speech to text with just one click. In addition, this time press and hold the input box (+1), press and hold to speak in voice mode (+1), as well as the voice key of the input method (+1), and the system's own dictation (+1).
Now when you open WeChat, there are 5 voice portals in the lower half of the screen waiting for you to speak.

Guigui, do you really want to listen to Mr. Bad Review’s subwoofer?
It would be fine if this situation only appeared in WeChat on the mobile phone. The problem is that many apps are like this now, from the mobile phone to the computer.
For example, the Mac desktop version of WeChat had the same prompt before, telling you to press and hold Fn to enter text by voice.

In the input box on the Feishu desktop, there is also a line "Hold down Fn to speak" to convert speech to text.

Brothers, you really don’t know how important Fn is on Mac, right? Are all the high-frequency keys in the input method and functional area occupied by your voice?
Don't worry, this is where we are going.
In the Feishu document, if you open a grid in a table, you will find that there is also a speech-to-text text in the upper right corner.

Let’s put it this way, if you open any big factory app now, you will see a bunch of microphones waving to you: Uncle, come to me...

Let’s be reasonable and make a voice input, no one has any objections, multiple choices are always a good thing.
But the current posture is obviously more than "giving you more options":
Some use text to remind you, some put the entrance in the most conspicuous place, and some simply make the button very large.
It seems that the major manufacturers have made an agreement to prioritize voice, for fear that we will not speak in the app.
Seeing this, you must be curious why, what is the magic of this voice input?
At first, the bad reviewer thought it was quite simple, it was easy to use and cheap.
The input efficiency is high compared to Pinyin, and the cost is not high.
According to the price on the official website of Volcano Engine, Doubao’s streaming speech recognition costs less than 1 yuan per hour. Even if a heavy user sends 10 minutes of speech every day, it only costs 60 yuan a year. This is still the retail price, and it will only be cheaper for major manufacturers to use their own models.

But the problem is:
Voice input is not new at all. Siri has been able to take dictation more than ten years ago, and voice search on Amap has been available for many years. Why is voice priority so high now?
Later, I thought about it and realized that the key may lie in the input method.
I have written before about the input methods of major manufacturers, saying that the input method is like a security guard guarding the door of all apps. It knows all your needs first, and then the app.
For example, if you ask a question in the Doubao App, the WeChat input method may have thrown the answer in your face before you click send.

Similarly, if you send a message in WeChat and type using Beanbao input method, WeChat definitely doesn’t want to see this scene.
Many years ago, Sogou input method surpassed Baidu.
The users are all on Baidu, and the candidate words that appear when typing are directly taken to Sogou for search.

So Apps must be thinking: Is there a way for users to interact directly with me and bypass the security of the input method?
Speech to text.
Indeed, this technology existed many years ago, but at that time the cost was high and the accuracy was low.
With the development of large-scale speech models in the past two years, its accuracy and response speed have become almost perfect.
It can be said that the voice-to-text supported by AI gives the app the opportunity to work alone for the first time without fear of being intercepted by the input method.

But this is just "defense". Let’s look a little further: major manufacturers are actually on the offensive, seizing the entrance to the next generation of interactions in advance.
Today, if you press and hold to speak, the App will help you convert your words into words; tomorrow, if you press and hold to speak, will the App directly book a flight for you?
There are already some AI-native apps that can do this, but apps like WeChat and Feishu also want to go in this direction, so they first train users to speak, and then train users to have direct conversations with Agents to complete their needs.
In the next generation of apps, you must open your mouth and speak → AI will do the work directly.
But this process needs to be gradual.
So now major apps must try their best to guide and cultivate the habit of users speaking in their own apps, and then gradually integrate various AI summary and search functions into voice commands.
This is why everyone is pushing voice recognition.

But this is not the end.
This phenomenon will also spread to most apps on the market in the future.
Because in the incremental era, everyone competes with who can run faster, while in the stock era, everyone competes with who can bite harder. When the cake no longer gets bigger, each family can only grab what is in their own bowl.
In this situation, any function that can be combined with AI and make users dependent, as long as the first company does it, the second company will follow. Because users are dependent, market demand will change, and if you don’t do it, you will be robbed of users.
Don’t believe me, just criticize yourself. I am a living example of being successfully trained by a big company.
When I first saw these voice entries, I thought to myself: I have been typing for more than ten years and my hand speed has not been slow, so why should I talk to my phone?
After trying it several times, I found that the recognition is really fast and accurate, and any ambiguities can be automatically corrected.
A few months ago I still defaulted to typing, and I just used voice because I was too lazy to do it...
What now?
I never type when I can open my mouth. Even when I am sitting in the office writing a manuscript, I want to dictate. My fingers are no longer sore and my input speed is faster. On the contrary, when I occasionally switch back to the input method, I wonder why it is so strenuous.

Of course my dependencies on AI go far beyond this.
Now that I see a long article, I don’t have the patience to read it from beginning to end, so I just leave it to the AI to summarize the key points. If I want to verify any news or videos on Twitter, @Grok will let it go; when I usually search for information, I simply read the AI summary, and if it’s difficult, I’ll let various agents sort it out for me.

I didn’t take much notice of these functions when they first appeared, but after experiencing them a few times, I couldn’t live without them without even realizing it.
When you and I both rely on AI like this, the market demand is subtly rewritten.
Everyone will not be able to use an App without AI voice and AI summarization. In order to adapt to the market, App developers will naturally have to follow suit to provide these functions.
Perhaps soon, the input method will no longer be a necessity.
When we first started talking about this topic, the editorial team was still joking:
If voice input continues like this, young people in the future may not even be able to type pinyin.
This may sound a bit scary at first, but when the day comes, it won’t matter if you can’t type in pinyin? Each generation has its own skills. The previous generation thought we couldn’t repair light bulbs or type five-stroke typing, but we still lived a pretty good life.
Input method has never been the goal, it is just a way to connect people and machines. When one day a new road becomes more convenient, no one will miss the old road.
The wheel of the times is rolling forward, and there are always some skills that will slowly disappear with the evolution of the way society operates.
Comments