On the future of voice interfaces, a rough, over-caffeinated brain dump:
ChatGPT’s advanced voice mode showed us what’s possible with interactivity. Notebook LM’s audio summaries showed us what’s possible with grounded conversational audio generation.
Both are incredible accomplishments and useful tools for thought.
What’s still missing? Let’s time travel in our DeLorean and see what the future may hold…
- Auto-generation. Today, sources need manual addition to Notebook LM, and audio needs regeneration. What if the product did this automatically? E.g., read my newsletter subscriptions or Slack messages from the last 24 hours and summarize them in audio. Additional sources: YouTube playlists, Notion folders, etc. Subject to constrains and customizations.
- Customization. While both products accept custom instructions, they don’t strictly follow them. This makes it hard to optimize content for specific use cases. Users will likely choose different settings for different use cases. Likely, smart UI defaults/knobs would be helpful in addition to text/voice instructions. Potential dimensions:
- Length
- Tone and format
- Number of speakers
- Abstraction/summarization level
- Educational density vs. entertainment/engagement
- Strict source adherence vs. source-inspired exploration
- Interactivity. It would be powerful to marry interactivity (advanced voice mode) with high-quality generated content (audio summaries). Imagine joining a pre-generated podcast about any topic as a third participant—either listening silently or interrupting to ask questions and redirect the conversation. This would make many GPUs go brr to achieve sufficiently low latency.
Some hallucination would still be acceptable for these use cases, so we’ll probably build them before we build truly reliable “nine nines” agentic behavior that is high-stakes.
Notes:
- For those unfamiliar: advanced voice mode enables natural conversation with the LLM, and Notebook LM audio summaries create AI-generated “podcasts” about anything.