Evom Research · August 2026

Đánh giá hiệu năng Loli và Loly trên độ chính xác, độ tự nhiên, độ trễ và khả năng xử lý tiếng Việt trong môi trường thực tế.

Loli · Speech-to-TextLoly · Text-to-Speech

3.8%

Vietnamese WER

110 ms

Loly TTFB

4.45 / 5

Naturalness MOS

~680 ms

End-to-end latency

Scroll to explore

Benchmark overview

Vietnamese-first. Real-time by design.

Các benchmark được thực hiện trong cùng môi trường để đo cả chất lượng model lẫn hiệu năng khi triển khai production.

01

Vietnamese-first Speech AI

Tối ưu cho giọng Bắc, Trung, Nam; hội thoại tự nhiên, môi trường nhiễu, tên riêng, địa chỉ, số tiền và thuật ngữ chuyên ngành.

02

Real-time Performance

Được thiết kế cho Voice Agent, Contact Center, virtual assistant, robotics, real-time transcription và smart devices.

03

Enterprise Ready

Đo concurrency, GPU utilization, P50/P95/P99, streaming throughput và độ ổn định khi vận hành dài hạn.

Loli 2.0 · Speech-to-Text

Listen with accuracy. Respond without delay.

Overall performance
Vietnamese WER3.8%< 5%
Vietnamese CER1.6%< 2.5%
First Partial Latency94 ms< 200 ms
Final Transcript Latency260 ms< 500 ms
Real-Time Factor0.08< 0.10
StreamingSupportedReal-time
Speaker DiarizationSupportedMulti-speaker

Vietnamese accent benchmark · WER

Northern Vietnamese2.9%
Central Vietnamese5.2%
Southern Vietnamese3.5%
Overall Vietnamese3.8%

Đánh giá riêng biệt trên ba vùng giọng thay vì chỉ sử dụng một tập tiếng Việt tổng hợp.

Difficult speech recognition tests

Built beyond studio audio.

Loli được kiểm thử trên các tình huống thường xuất hiện trong cuộc gọi và môi trường thực tế.

Clean SpeechBackground NoiseCall Center 8 kHzFast SpeechQuiet SpeechMultiple SpeakersVietnamese + EnglishProper NamesAddressesNumbersCurrencyDates & TimeProduct NamesIndustry Terminology

Critical information recognition

Every number changes the meaning.

“Tài khoản của quý khách vừa phát sinh giao dịch 18 triệu 500 nghìn đồng lúc 14 giờ 35 phút.”
18.500.000 VND14:35IntentCurrencyNormalization

Real-time STT latency

User SpeechAudio stream
Voice Activity DetectionVAD
First Partial Transcript94 ms
Final Transcript260 ms
Real-Time Factor0.08

Loly 3.5 · Text-to-Speech

Natural voice. First audio in 110 milliseconds.

Overall performance
TTFB110 ms< 200 ms
Real-Time Factor0.06< 0.10
Naturalness MOS4.45 / 5> 4.0
Speech Clarity4.61 / 5> 4.0
Emotion Score4.36 / 5> 4.0
Speaker Similarity0.92> 0.85
Streaming TTSSupportedReal-time

Mean Opinion Score · MOS

Naturalness4.45/ 5
Speech Clarity4.61/ 5
Emotion & Prosody4.36/ 5

Blind evaluation by independent listeners on a five-point scale.

Voice cloning benchmark

Speaker similarity

0.92
01Reference Voice02Speaker Embedding03Generated Voice04Cosine Similarity

Streaming TTS performance

Time to first audio

110 ms
01Request · 0 ms02First Audio Chunk · 110 ms03Continuous Streaming04Generation Complete

Voice Agent end-to-end

From speech to response in ~680 ms.

Speech AI được đánh giá như một hệ thống hoàn chỉnh, từ lúc người dùng dừng nói đến khi nghe thấy audio đầu tiên từ AI.

01

VAD

100 ms

02

Loli STT

180 ms

03

LLM First Token

200 ms

04

Loly TTS

120 ms

05

Network

80 ms

Loli STT latency distribution

Latency
P50
180 ms
P95
290 ms
P99
420 ms

Loly TTS latency distribution

TTFB
P50
110 ms
P95
180 ms
P99
290 ms

Vietnamese speech benchmark dataset

1,000+ sentences. Real-world Vietnamese.

Bộ kiểm thử được chia thành nhiều nhóm dữ liệu để phản ánh điều kiện sử dụng thực tế tại Việt Nam.

01

General Conversation

Hội thoại đời sống và giao tiếp tự nhiên.

02

Customer Service

Tư vấn, khiếu nại và chăm sóc khách hàng.

03

Banking & Finance

Giao dịch, tài khoản, số tiền và xác thực.

04

Insurance

Hợp đồng, quyền lợi và tư vấn bảo hiểm.

05

Healthcare

Thuật ngữ y tế, lịch khám và hướng dẫn bệnh nhân.

06

E-commerce

Sản phẩm, đơn hàng, vận chuyển và thanh toán.

07

Numbers & Dates

Số điện thoại, OTP, ngày giờ và số tiền.

08

Proper Nouns

Tên người, công ty, thương hiệu và địa danh.

09

Regional Accents

Giọng Bắc, Trung và Nam.

10

Difficult Conditions

Tiếng ồn, mic chất lượng thấp và tốc độ nói cao.

Benchmark Comparison

Loli & Loly vs leading Speech AI models.

We evaluate every platform with the same Vietnamese-focused methodology—measuring real-world Voice Agents, Contact Centers, Banking, Healthcare, E-commerce and Robotics rather than clean studio audio alone.

Loli & Loly
ElevenLabs
OpenAI
Deepgram
Google Cloud
Microsoft Azure

Speech-to-Text Comparison

Loli vs leading STT models.

Vietnamese accuracy is measured separately from global speech scores. Lower WER and CER are better.

ModelVietnamese WER ↓Vietnamese CER ↓Streaming
Loli 2.03.8%*1.6%*
Scribe v24.4%1.9%
OpenAI Speech-to-Text4.7%2.1%
Nova-35.2%2.4%
Google Cloud STT5.6%2.6%
Azure Speech5.9%2.8%

* Preliminary Evom Labs internal benchmark. Final results will be published after validation on the complete benchmark dataset.

Why we benchmark Vietnamese separately

Global speech benchmarks do not necessarily reflect performance in Vietnam. The dataset explicitly covers the way Vietnamese people speak across regions, industries and channels.

Northern VietnameseCentral VietnameseSouthern VietnameseVietnamese-English mixed speechLocal names & addressesPhone numbers & currencyBanking & healthcare terminologyCall-center and 8 kHz audio

Regional Accent Accuracy

Understanding Vietnam means understanding its accents.

Instead of hiding regional differences inside a single average score, Evom Labs measures every major Vietnamese accent independently.

ModelNorth WER ↓Central WER ↓South WER ↓
Loli 2.02.9%*5.2%*3.5%*
ElevenLabs ScribeTest PendingTest PendingTest Pending
OpenAITest PendingTest PendingTest Pending
Deepgram Nova-3Test PendingTest PendingTest Pending
GoogleTest PendingTest PendingTest Pending
AzureTest PendingTest PendingTest Pending

Real-World Vietnamese STT

Beyond clean studio audio.

Production audio includes noise, low-bandwidth telephony, overlapping speakers and critical local entities.

Scenario / CapabilityLoli & LolyElevenLabsOpenAIDeepgramGoogle CloudMicrosoft Azure
Clean Vietnamese
Northern Accent
Central Accent
Southern Accent
8 kHz Call Center
Background Noise
Multiple Speakers
Vietnamese + English
Local Proper NamesOptimizedGeneralGeneralGeneralGeneralGeneral
Vietnamese CurrencyOptimizedGeneralGeneralGeneralGeneralGeneral
Vietnamese AddressesOptimizedGeneralGeneralGeneralGeneralGeneral

* “Optimized” indicates an explicit focus in Loli evaluation and development. It is not an independently verified superiority claim until the comparative benchmark is complete.

Critical Information Recognition

Numbers matter.

For Voice AI, recognizing the sentence is not enough. A dedicated score evaluates numbers, currency, phone numbers, names and addresses.

18.500.000 đồng
09:45
OTP 837291
0987 654 321
Nguyễn Văn Thành
Hoàng Quốc Việt, Hà Nội
ModelNumbersCurrencyPhoneNamesAddresses
Loli & LolyTBDTBDTBDTBDTBD
ElevenLabsTBDTBDTBDTBDTBD
OpenAITBDTBDTBDTBDTBD
DeepgramTBDTBDTBDTBDTBD
Google CloudTBDTBDTBDTBDTBD
Microsoft AzureTBDTBDTBDTBDTBD
BankingInsuranceHealthcareCustomer ServiceLogisticsE-commerceGovernment Services

STT Speed Comparison

Accuracy without sacrificing speed.

All final latency measurements use the same client location, network and test harness.

ModelFirst Partial ↓Finalization ↓Real-Time
Loli 2.094 ms*260 ms*
Scribe RealtimeBenchmark PendingBenchmark Pending
Nova-3Benchmark PendingBenchmark Pending
OpenAIBenchmark PendingBenchmark Pending
GoogleBenchmark PendingBenchmark Pending
AzureBenchmark PendingBenchmark Pending

Independent context: Deepgram Nova-3

Deepgram publicly reports 6.84% median streaming WER and 5.26% batch WER on its own multi-domain benchmark. Different datasets mean these figures are useful context—not a direct comparison with Loli's Vietnamese WER.

Independent context: ElevenLabs Scribe

ElevenLabs positions Scribe v2 for high-accuracy batch transcription and Scribe v2 Realtime for low-latency agents. Evom Labs re-tests it on the same Vietnamese dataset for a fair comparison.

Text-to-Speech Comparison

Loly vs leading TTS models.

The primary real-time metric is TTFA—Time To First Audio—because users care when the AI starts speaking.

ModelFirst Audio ↓Streaming
Loly 3.5110 ms*
ElevenLabs Flash v2.5~75 ms inference†
OpenAITo be measured
GoogleTo be measured
AzureTo be measured

A fair latency comparison

Model latency ≠ user-perceived latency.

01Network
02Queue
03Inference
04Streaming buffer
05Playback

Final comparison measures Request Sent → First Audio Received → First Audio Played, rather than mixing provider marketing numbers with end-to-end measurements.

Loly vs ElevenLabs Flash

Two real-time references. One fair test.

Flash v2.5 is a primary low-latency reference. Its reported ~75 ms model inference excludes network and application overhead, so it is not equivalent to end-to-end TTFA.

MetricLoly 3.5ElevenLabs Flash v2.5
VietnameseNative focusSupported
Streaming
First Audio110 ms*~75 ms inference†
Vietnamese Number HandlingOptimized*Text normalization available
Voice Cloning
Real-time Voice Agent
On-Premise DeploymentAvailable*Depends on offering
Private DeploymentAvailable*Depends on offering

Naturalness Comparison

Blind human listening—not brand preference.

Listeners do not know which model produced each sample. Maximum MOS is 5.0; higher is better.

ModelNaturalness ↑Clarity ↑Emotion ↑
Loly 3.54.45*4.61*4.36*
ElevenLabsBenchmark PendingBenchmark PendingBenchmark Pending
OpenAIBenchmark PendingBenchmark PendingBenchmark Pending
GoogleBenchmark PendingBenchmark PendingBenchmark Pending
AzureBenchmark PendingBenchmark PendingBenchmark Pending

Designed to sound Vietnamese—not just speak Vietnamese

The pronunciation set targets phrases that are frequently difficult for multilingual systems.

Nguyễn Thị Thu HuyềnTrần Quốc KhánhBuôn Ma ThuộtĐắk Lắk18 triệu 500 nghìn đồng14 giờ 35 phútOpenAI APIAI Voice Agent
ModelVietnameseNumbersNamesCurrencyMixed EN-VI
Loli & LolyTBDTBDTBDTBDTBD
ElevenLabsTBDTBDTBDTBDTBD
OpenAITBDTBDTBDTBDTBD
Google CloudTBDTBDTBDTBDTBD
Microsoft AzureTBDTBDTBDTBDTBD

Voice Cloning Comparison

How close is the generated voice?

Machine speaker-embedding similarity is reported together with blind human preference testing.

ModelSpeaker Similarity ↑
Loly 3.50.92*
ElevenLabsBenchmark Pending
OpenAIN/A / capability dependent
GoogleCapability dependent
AzureCapability dependent
TimbrePitchSpeaking rhythmAccentPronunciationSpeaking styleEmotion

Voice Agent Benchmark

What actually matters: end-to-end latency.

From the moment the user finishes speaking until the first AI audio begins.

01

End-of-Speech

100 ms
02

Loli STT

180 ms
03

LLM TTFT

200 ms
04

Loly TTFA

120 ms
05

Network

80 ms

End-to-End Response

~680 ms*

Preliminary internal benchmark—from user speech completion to first AI audio.

Voice AI Platform Comparison

Beyond model quality.

Enterprise deployments also depend on privacy, customization and where inference can run.

Scenario / CapabilityLoli & LolyElevenLabsOpenAIDeepgramGoogle CloudMicrosoft Azure
STT
TTSLimited / API dependent✓ / product dependent
Streaming
Vietnamese FocusCoreMultilingualMultilingualMultilingualMultilingualMultilingual
Vietnamese Accent Benchmark
Voice CloningCapability dependentCapability dependentCapability dependent
Voice Agent Ready
On-PremiseOffering dependentPrimarily cloudOffering dependentCloud / enterpriseCloud / enterprise
Private VPCEnterprise dependentEnterprise dependentEnterprise dependent
Air-Gapped✓*Deployment dependentDeployment dependentDeployment dependent

Where Loli & Loly Are Different

Vietnamese-first by design.

One coordinated speech stack for how Vietnamese people actually communicate.

01

Vietnamese-First

Vietnamese sits at the center of model development and evaluation.

02

Regional Accents

North, Central and South are benchmarked independently.

03

Critical Information

Numbers, currency, names, addresses and mixed speech get dedicated tests.

04

Real-Time Voice AI

Loli listens. AI understands. Loly speaks.

05

Enterprise Deployment

Private architectures for sensitive voice data and regulated workloads.

Benchmark Summary

Speech-to-Text

Loli 2.0

Vietnamese WER3.8%*
First Partial94 ms*
Final Transcript260 ms*

Text-to-Speech

Loly 3.5

TTFA110 ms*
Naturalness MOS4.45 / 5*
Speaker Similarity0.92*

Voice Agent

Loli + Loly

End-to-End~680 ms*
Language FocusVietnamese
ModeReal-time

Run the Benchmark Yourself

Don't take our numbers at face value.

Upload the same audio or enter the same Vietnamese text, then compare every model under equivalent conditions.

Compare STT

Compare transcript, WER, CER, processing time, latency, numbers, names and Vietnamese accents.

Upload Your Audio

Compare TTS

Listen blind and compare naturalness, pronunciation, emotion, voice similarity and first-audio latency.

Test Loly
Benchmark Disclosure. Competitor performance varies by model version, region, API configuration, network and deployment environment. External figures use different methodologies. Evom figures marked * are preliminary internal measurements and should be independently validated before use as public performance claims.

Benchmark methodology

Measured under the same conditions.

GPU

NVIDIA RTX A5000

Batch Size

1

Audio

16 kHz / Mono

Warm-up

10 requests

Measured Requests

500+

Concurrency

1 / 10 / 25 / 50 / 100

Metrics

P50 / P95 / P99

Evaluation

Multiple runs + blind MOS

Benchmark transparency. Kết quả có thể thay đổi theo phần cứng, cấu hình triển khai, điều kiện audio, network latency và phiên bản model. Các số liệu trên đang ở trạng thái benchmark preview và cần được xác nhận bằng kết quả chính thức trước khi dùng trong tài liệu cam kết thương mại.

Speech-to-Text

Loli

Listen.

Reasoning & Agent

AI

Understand.

Text-to-Speech

Loly

Speak.

Test Loli & Loly

Hear the benchmark. Test the models yourself.

Last update: August 2026 · Loli 2.0 · Loly 3.5