Samuel Zeng

29 posts

Samuel Zeng banner
Samuel Zeng

Samuel Zeng

@SamuelZengML

Founder, https://t.co/OuQHbxHus6 | Building the voice layer for AI: cloud speech foundation models + on-device intelligence | MIT TR35 https://t.co/WQ9mrGQA21

San Francisco Bay Area Katılım Ekim 2012
174 Takip Edilen285 Takipçiler
Sabitlenmiş Tweet
Samuel Zeng
Samuel Zeng@SamuelZengML·
The future of voice is hybrid. I'm building Audio8.ai around a stack: cloud speech foundation models plus on-device intelligence. Cloud brings quality and scale. Devices add privacy, low latency, and reliability. The goal: make voice native to every AI system.
English
0
0
2
754
Samuel Zeng
Samuel Zeng@SamuelZengML·
Audio8 TTS Preview 0.6B ONNX INT4 is now available. A compact model with serious capability: - Runs entirely on CPU with ONNX Runtime - Uses only around 1 GB of memory during inference - Multilingual TTS across 11 languages - Zero-shot voice cloning - No PyTorch or Transformers required - Includes CLI, streaming, HTTP/OpenAI-compatible APIs, and voice registration Small footprint. Full voice-cloning capability. Small footprint. Full voice-cloning capability. Model: huggingface.co/Audio8/Audio8-… Code: github.com/Audio8-AI/Audi…
English
11
32
288
11.6K
Samuel Zeng
Samuel Zeng@SamuelZengML·
Audio8 ASR is part of a broader thesis: high-quality speech intelligence should be open, multilingual, and deployable across cloud and edge. Paper: arxiv.org/abs/2605.28139 Try the models, star the repo, and tell us what you build.
English
0
0
4
529
Samuel Zeng
Samuel Zeng@SamuelZengML·
Training, inference, evaluation, and vLLM serving code are open. The repo includes Transformers inference, batch evaluation, a native /asr endpoint, and an OpenAI-compatible /v1/audio/transcriptions endpoint. GitHub: github.com/AutoArk/open-a…
English
2
0
1
595
Samuel Zeng
Samuel Zeng@SamuelZengML·
Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a 4.76 Mean WER. Our 0.6B model also ranks in the Top 5. Two model sizes. One open speech stack. Thread ↓
Samuel Zeng tweet media
English
11
28
201
31.9K
Samuel Zeng
Samuel Zeng@SamuelZengML·
Open source means more than a checkpoint: • weights + bundled neural codec • zero-shot/no-reference inference • JSONL batch generation • codec-index preparation • single-/multi-GPU SFT Build with it. Break it. Tell us where it fails. github.com/Audio8-AI/Audi… 4/4
Samuel Zeng tweet media
English
2
1
56
4K
Samuel Zeng
Samuel Zeng@SamuelZengML·
Benchmarks aren’t enough for speech. Hear it. The listening set covers multilingual speech, pronunciation challenges, tongue twisters, classical Chinese, hard cases, and cross-lingual voice cloning. Listen: audio8-ai.github.io/Audio8_TTS/ 3/4
Samuel Zeng tweet media
English
2
1
49
5.1K
Samuel Zeng
Samuel Zeng@SamuelZengML·
A 0.6B speech model should not sound this capable. Today, we’re open-sourcing Audio8-TTS Preview: 11 languages. Zero-shot voice cloning. A bundled 44.1 kHz codec. Apache 2.0. Small model. Serious speech. Thread ↓
Samuel Zeng tweet media
English
43
148
1.4K
108.5K
Samuel Zeng
Samuel Zeng@SamuelZengML·
For reproducibility, the model card reports a 7.03 WER mean across the seven current Open ASR Leaderboard splits. The dataset revision, inference setup, and raw manifests are documented here: huggingface.co/AutoArk-AI/Aud…
English
0
0
0
70
Samuel Zeng
Samuel Zeng@SamuelZengML·
We just open-sourced Audio8-ASR-0.1B, one of the smallest production-ready ASR models with cloud-level accuracy. Only 100M parameters, under 200MB peak memory, and built to run on real hardware. No cloud required. huggingface.co/AutoArk-AI/Aud…
Samuel Zeng tweet media
English
2
8
33
2.2K
Samuel Zeng
Samuel Zeng@SamuelZengML·
From meeting notes and live captions to voice assistants and smart hardware, we want high-quality ASR to run directly on every device, without being limited by cloud access or compute. Small model. SOTA-level performance. Fully on-device. GitHub: github.com/AutoArk/open-a… Paper: arxiv.org/abs/2605.28139
English
0
2
2
276
Samuel Zeng
Samuel Zeng@SamuelZengML·
Built for on-device deployment: ✅ ONNX Runtime + Apple Neural Engine ✅ Local transcription with low latency and high privacy ✅ 7 languages: Mandarin, English, French, German, Japanese, Korean, and Cantonese ✅ Hotword boosting without fine-tuning
English
0
0
2
468
Samuel Zeng retweetledi
Musolsol.𝟎𝐱𝐔
Musolsol.𝟎𝐱𝐔@MMMusol·
网易云花了十年都没做出来的歌词体验, 被一个独立开发者用开源项目实现了 Folia —— 音乐可视化+动画歌词,丝滑程度头皮发麻 不夸张地说,这是我见过审美最在线的音乐类开源项目,没有之一,UI 细节拉满 强烈推荐给所有对音乐播放体验和审美有追求的朋友 GitHub:github.com/chthollyphile/…
Musolsol.𝟎𝐱𝐔@MMMusol

艺术性和视觉化夯爆了 真的是近期见过最惊艳的音乐播放器了

中文
75
120
723
121.9K
Samuel Zeng
Samuel Zeng@SamuelZengML·
🚀 Just open-sourced OpenResearch! AutoResearch + AlphaEvolve-style local search for code, prompts & training loops. Key features: • Mission-driven autonomous iteration • Persistent memory + fully replayable artifacts • Adapters for prompt eval, benchmark tuning & real GPU training • Beautiful HTML reports & proposal journals • 100% local, uv one-click, works with any OpenAI-compatible LLM → github.com/SamuelZeng/Ope… Star ⭐, try it, contribute! Let's automate research together 🔥 #OpenSource #AI #LLM #AutonomousAgents #MachineLearning
English
1
0
2
138