這種轉變已經開始改變使用者的行為模式。Google指出,其即時語音服務的使用量在過去一年內增加了一倍,且語音對話的長度平均是文字對話的五倍。OpenAI也表示每週有超過一億五千萬人使用ChatGPT的語音和聽寫功能。許多使用者在進行寫程式等複雜任務時,發現用語音更能完整表達自己的意圖、限制與期望的結果。(關鍵數字:150)
儘管語音技術前景看好,但仍面臨背景噪音干擾以及在公共場合對著機器說話的社交尷尬等挑戰。為了解決這些痛點,矽谷的新創公司紛紛投入開發更隱密且可靠的語音輸入硬體設備,例如特殊麥克風和智慧戒指等。隨著相關創投金額大幅成長,語音互動有望在未來克服技術障礙,成為最主流的操作介面。

Tech companies are actively working to make AI voice assistants sound more natural, betting that voice will replace typing as the primary mode of interaction. Both OpenAI and Google have recently upgraded their voice models by processing audio in real time instead of converting it to text first. This advancement significantly reduces response latency, making the AI sound less like a robotic reader and more like an intuitive colleague.
This shift in technology is already influencing user behavior in noticeable ways. Google reported that its live voice usage has doubled over the past year, with voice conversations lasting five times longer than text-based ones. Similarly, OpenAI sees over 150 million weekly users utilizing ChatGPT's voice features. Users are finding that speaking allows them to articulate complex intentions, constraints, and desired outcomes more effectively, particularly in tasks like coding.
Despite the promising future of voice technology, it still faces hurdles such as background noise interference and the social discomfort of talking to machines in public. To address these issues, Silicon Valley start-ups are developing discreet and reliable hardware solutions, including specialized microphones and smart rings. With venture capital investment in this sector surging, voice interaction is poised to overcome these technical barriers and become the dominant paradigm.