Vietnamese-first Speech AI
Tối ưu cho giọng Bắc, Trung, Nam; hội thoại tự nhiên, môi trường nhiễu, tên riêng, địa chỉ, số tiền và thuật ngữ chuyên ngành.
Evom Research · August 2026
Đánh giá hiệu năng Loli và Loly trên độ chính xác, độ tự nhiên, độ trễ và khả năng xử lý tiếng Việt trong môi trường thực tế.
3.8%
Vietnamese WER
110 ms
Loly TTFB
4.45 / 5
Naturalness MOS
~680 ms
End-to-end latency
Benchmark overview
Các benchmark được thực hiện trong cùng môi trường để đo cả chất lượng model lẫn hiệu năng khi triển khai production.
Tối ưu cho giọng Bắc, Trung, Nam; hội thoại tự nhiên, môi trường nhiễu, tên riêng, địa chỉ, số tiền và thuật ngữ chuyên ngành.
Được thiết kế cho Voice Agent, Contact Center, virtual assistant, robotics, real-time transcription và smart devices.
Đo concurrency, GPU utilization, P50/P95/P99, streaming throughput và độ ổn định khi vận hành dài hạn.
Loli 2.0 · Speech-to-Text
Vietnamese accent benchmark · WER
Đánh giá riêng biệt trên ba vùng giọng thay vì chỉ sử dụng một tập tiếng Việt tổng hợp.
Difficult speech recognition tests
Loli được kiểm thử trên các tình huống thường xuất hiện trong cuộc gọi và môi trường thực tế.
Critical information recognition
“Tài khoản của quý khách vừa phát sinh giao dịch 18 triệu 500 nghìn đồng lúc 14 giờ 35 phút.”
Real-time STT latency
Loly 3.5 · Text-to-Speech
Mean Opinion Score · MOS
Blind evaluation by independent listeners on a five-point scale.
Voice cloning benchmark
Streaming TTS performance
Voice Agent end-to-end
Speech AI được đánh giá như một hệ thống hoàn chỉnh, từ lúc người dùng dừng nói đến khi nghe thấy audio đầu tiên từ AI.
100 ms
180 ms
200 ms
120 ms
80 ms
Vietnamese speech benchmark dataset
Bộ kiểm thử được chia thành nhiều nhóm dữ liệu để phản ánh điều kiện sử dụng thực tế tại Việt Nam.
Hội thoại đời sống và giao tiếp tự nhiên.
Tư vấn, khiếu nại và chăm sóc khách hàng.
Giao dịch, tài khoản, số tiền và xác thực.
Hợp đồng, quyền lợi và tư vấn bảo hiểm.
Thuật ngữ y tế, lịch khám và hướng dẫn bệnh nhân.
Sản phẩm, đơn hàng, vận chuyển và thanh toán.
Số điện thoại, OTP, ngày giờ và số tiền.
Tên người, công ty, thương hiệu và địa danh.
Giọng Bắc, Trung và Nam.
Tiếng ồn, mic chất lượng thấp và tốc độ nói cao.
Benchmark Comparison
We evaluate every platform with the same Vietnamese-focused methodology—measuring real-world Voice Agents, Contact Centers, Banking, Healthcare, E-commerce and Robotics rather than clean studio audio alone.
Speech-to-Text Comparison
Vietnamese accuracy is measured separately from global speech scores. Lower WER and CER are better.
| Model | Vietnamese WER ↓ | Vietnamese CER ↓ | Streaming |
|---|---|---|---|
| 3.8%* | 1.6%* | ✓ | |
| 4.4% | 1.9% | ✓ | |
| 4.7% | 2.1% | ✓ | |
| 5.2% | 2.4% | ✓ | |
| 5.6% | 2.6% | ✓ | |
| 5.9% | 2.8% | ✓ |
* Preliminary Evom Labs internal benchmark. Final results will be published after validation on the complete benchmark dataset.
Global speech benchmarks do not necessarily reflect performance in Vietnam. The dataset explicitly covers the way Vietnamese people speak across regions, industries and channels.
Regional Accent Accuracy
Instead of hiding regional differences inside a single average score, Evom Labs measures every major Vietnamese accent independently.
| Model | North WER ↓ | Central WER ↓ | South WER ↓ |
|---|---|---|---|
| 2.9%* | 5.2%* | 3.5%* | |
| Test Pending | Test Pending | Test Pending | |
| Test Pending | Test Pending | Test Pending | |
| Test Pending | Test Pending | Test Pending | |
| Test Pending | Test Pending | Test Pending | |
| Test Pending | Test Pending | Test Pending |
Real-World Vietnamese STT
Production audio includes noise, low-bandwidth telephony, overlapping speakers and critical local entities.
| Scenario / Capability | ||||||
|---|---|---|---|---|---|---|
| Clean Vietnamese | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Northern Accent | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Central Accent | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Southern Accent | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 8 kHz Call Center | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Background Noise | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Multiple Speakers | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Vietnamese + English | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Local Proper Names | Optimized | General | General | General | General | General |
| Vietnamese Currency | Optimized | General | General | General | General | General |
| Vietnamese Addresses | Optimized | General | General | General | General | General |
* “Optimized” indicates an explicit focus in Loli evaluation and development. It is not an independently verified superiority claim until the comparative benchmark is complete.
Critical Information Recognition
For Voice AI, recognizing the sentence is not enough. A dedicated score evaluates numbers, currency, phone numbers, names and addresses.
| Model | Numbers | Currency | Phone | Names | Addresses |
|---|---|---|---|---|---|
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD |
STT Speed Comparison
All final latency measurements use the same client location, network and test harness.
| Model | First Partial ↓ | Finalization ↓ | Real-Time |
|---|---|---|---|
| 94 ms* | 260 ms* | ✓ | |
| Benchmark Pending | Benchmark Pending | ✓ | |
| Benchmark Pending | Benchmark Pending | ✓ | |
| Benchmark Pending | Benchmark Pending | ✓ | |
| Benchmark Pending | Benchmark Pending | ✓ | |
| Benchmark Pending | Benchmark Pending | ✓ |
Deepgram publicly reports 6.84% median streaming WER and 5.26% batch WER on its own multi-domain benchmark. Different datasets mean these figures are useful context—not a direct comparison with Loli's Vietnamese WER.
ElevenLabs positions Scribe v2 for high-accuracy batch transcription and Scribe v2 Realtime for low-latency agents. Evom Labs re-tests it on the same Vietnamese dataset for a fair comparison.
Text-to-Speech Comparison
The primary real-time metric is TTFA—Time To First Audio—because users care when the AI starts speaking.
| Model | First Audio ↓ | Streaming |
|---|---|---|
| 110 ms* | ✓ | |
| ~75 ms inference† | ✓ | |
| To be measured | ✓ | |
| To be measured | ✓ | |
| To be measured | ✓ |
A fair latency comparison
Final comparison measures Request Sent → First Audio Received → First Audio Played, rather than mixing provider marketing numbers with end-to-end measurements.
Loly vs ElevenLabs Flash
Flash v2.5 is a primary low-latency reference. Its reported ~75 ms model inference excludes network and application overhead, so it is not equivalent to end-to-end TTFA.
| Metric | Loly 3.5 | ElevenLabs Flash v2.5 |
|---|---|---|
| Vietnamese | Native focus | Supported |
| Streaming | ✓ | ✓ |
| First Audio | 110 ms* | ~75 ms inference† |
| Vietnamese Number Handling | Optimized* | Text normalization available |
| Voice Cloning | ✓ | ✓ |
| Real-time Voice Agent | ✓ | ✓ |
| On-Premise Deployment | Available* | Depends on offering |
| Private Deployment | Available* | Depends on offering |
Naturalness Comparison
Listeners do not know which model produced each sample. Maximum MOS is 5.0; higher is better.
| Model | Naturalness ↑ | Clarity ↑ | Emotion ↑ |
|---|---|---|---|
| 4.45* | 4.61* | 4.36* | |
| Benchmark Pending | Benchmark Pending | Benchmark Pending | |
| Benchmark Pending | Benchmark Pending | Benchmark Pending | |
| Benchmark Pending | Benchmark Pending | Benchmark Pending | |
| Benchmark Pending | Benchmark Pending | Benchmark Pending |
The pronunciation set targets phrases that are frequently difficult for multilingual systems.
| Model | Vietnamese | Numbers | Names | Currency | Mixed EN-VI |
|---|---|---|---|---|---|
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD | |
| TBD | TBD | TBD | TBD | TBD |
Voice Cloning Comparison
Machine speaker-embedding similarity is reported together with blind human preference testing.
| Model | Speaker Similarity ↑ |
|---|---|
| 0.92* | |
| Benchmark Pending | |
| N/A / capability dependent | |
| Capability dependent | |
| Capability dependent |
Voice Agent Benchmark
From the moment the user finishes speaking until the first AI audio begins.
End-of-Speech
100 msLoli STT
180 msLLM TTFT
200 msLoly TTFA
120 msNetwork
80 msEnd-to-End Response
~680 ms*Preliminary internal benchmark—from user speech completion to first AI audio.
Voice AI Platform Comparison
Enterprise deployments also depend on privacy, customization and where inference can run.
| Scenario / Capability | ||||||
|---|---|---|---|---|---|---|
| STT | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| TTS | ✓ | ✓ | ✓ | Limited / API dependent | ✓ / product dependent | ✓ |
| Streaming | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Vietnamese Focus | Core | Multilingual | Multilingual | Multilingual | Multilingual | Multilingual |
| Vietnamese Accent Benchmark | ✓ | — | — | — | — | — |
| Voice Cloning | ✓ | ✓ | Capability dependent | — | Capability dependent | Capability dependent |
| Voice Agent Ready | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| On-Premise | ✓ | Offering dependent | Primarily cloud | Offering dependent | Cloud / enterprise | Cloud / enterprise |
| Private VPC | ✓ | Enterprise dependent | Enterprise dependent | Enterprise dependent | ✓ | ✓ |
| Air-Gapped | ✓* | Deployment dependent | — | Deployment dependent | — | Deployment dependent |
Where Loli & Loly Are Different
One coordinated speech stack for how Vietnamese people actually communicate.
Vietnamese sits at the center of model development and evaluation.
North, Central and South are benchmarked independently.
Numbers, currency, names, addresses and mixed speech get dedicated tests.
Loli listens. AI understands. Loly speaks.
Private architectures for sensitive voice data and regulated workloads.
Benchmark Summary
Speech-to-Text
Text-to-Speech
Voice Agent
Run the Benchmark Yourself
Upload the same audio or enter the same Vietnamese text, then compare every model under equivalent conditions.
Compare transcript, WER, CER, processing time, latency, numbers, names and Vietnamese accents.
Upload Your AudioListen blind and compare naturalness, pronunciation, emotion, voice similarity and first-audio latency.
Test LolyBenchmark methodology
NVIDIA RTX A5000
1
16 kHz / Mono
10 requests
500+
1 / 10 / 25 / 50 / 100
P50 / P95 / P99
Multiple runs + blind MOS
Speech-to-Text
Listen.
Reasoning & Agent
Understand.
Text-to-Speech
Speak.
Test Loli & Loly
Last update: August 2026 · Loli 2.0 · Loly 3.5