
I’d love to see future Gemma models move toward native TTS and truly real-time voice interaction.
I’d also be excited to see a broader range of small, efficient MoE models built for local and edge AI, alongside stronger real-time multimodality, more memory-efficient long-context inference, and more capable agents with better tool use.
Dynamic compute allocation based on task complexity would be a particularly interesting direction as well.
English





