TTS Model Distillation: From Cloud to Edge on RK3308

June 2026 | MetoClaw IoT Technology

Fast Voice Cloning: A Complete Guide to TTS Model Distillation

Tags: TTS Distillation · RK3308 · Qwen · MOS 4.5 · Embedded AI · CosyVoice
June 2026 · ~18 min read


1. The Next Frontier in Voice Interaction: From "Can Speak" to "Speaks Well"

Fig 1: TTS Pipeline Overview

If you bought a smart speaker in 2025, you've probably experienced this: it suddenly blurts out a robotic-sounding "Okay, turning on the living room light"—you can hear every syllable boundary but not a trace of human warmth. This isn't a hardware problem. It's the longstanding pain point of on-device TTS (Text-to-Speech): audio quality and compute power have always been a zero-sum trade-off.

That is now changing.

Speech large models led by Qwen have already reached near-human levels in the cloud—CosyVoice2, with just 0.5B parameters, achieves a Chinese CER as low as 1.45% and speaker similarity of 75.7%, making synthetic speech indistinguishable from real human voices to the average listener. The problem: these models need to run on A100/H800 GPUs, worlds apart from embedded hardware.

Meanwhile, the RK3308—an IoT voice chip priced under $3—has just four Cortex-A35 cores, 256MB DDR3 memory, and not even an NPU.

This article answers a seemingly impossible question: how to transfer the "intelligence" of Qwen's large speech model, through the art of distillation, onto the RK3308's tiny silicon, achieving TTS synthesis quality above MOS 4.5 in both Chinese and English.

MOS (Mean Opinion Score): The gold standard for speech quality, scored on a 1–5 scale. 4.0 = good (occasional perceptible artifacts), 4.5 = excellent (near professional recording quality), 5.0 = perfect (indistinguishable from a real human). Most embedded TTS systems on the market hover in the 3.0–3.8 range.

2. RK3308 Deep Dive: Dancing on a Pinhead

Fig 2: RK3308 Architecture

Figure: RK3308 internal architecture — Audio Codec and VAD are its core differentiated capabilities

2.1 Hardware Specifications at a Glance

Fig 3: Teacher Model Comparison

Parameter Specification Impact on TTS
CPU 4× Cortex-A35 @ 1.3GHz CPU-only inference; no GPU/NPU acceleration
RAM 256MB DDR3/DDR3L (max 512MB) Model + runtime must stay within 64MB
Storage SPI Nor/Nand Flash Model files require compressed storage
Audio 8ch ADC + 2ch DAC, 24bit/192kHz High-fidelity output, supports multi-mic arrays
VAD Hardware Voice Activity Detection Ultra-low-power wake-up, synergizes with TTS
Interfaces I2S/PCM/TDM, UART, SPI, USB Flexible external Codec/DSP connectivity
Power ~0.8–1.2W during TTS inference Battery-operable; suitable for offline devices
OS Linux (Buildroot/Yocto) Full ONNX Runtime support

2.2 Compute Budget: How Much Do You Get?

Fig 4: Knowledge Distillation Flow

The RK3308's single Cortex-A35 core delivers approximately 1.9 DMIPS/MHz, totaling roughly 10,000 DMIPS across all four cores. For reference, a single Cortex-A72 core on the Raspberry Pi 4 delivers around 25,000 DMIPS.

Translated to the TTS task:

2.3 The RK3308's "Secret Weapon": Hardware VAD

Fig 5: VITS-Mini Architecture

The RK3308's built-in hardware VAD (Voice Activity Detection) unit can continuously listen for human speech at under 1mW of power. Upon detecting voice, it wakes the CPU to start TTS inference. This provides the critical foundation for "always-on" voice interaction scenarios.

Key Insight: The RK3308's real value isn't in raw compute power—it lies in its end-to-end hardware audio pipeline: ADC capture → VAD wake-up → TTS synthesis → DAC output, all on a single chip with a minimal BOM cost.

3. Teacher Model Selection: The Qwen Speech Model Landscape

Fig 6: Deployment Pipeline

Figure: Capability matrix of the Qwen/CosyVoice model family

3.1 Why Qwen?

Fig 7: Bilingual Results

The Qwen series is an open-source large language model family developed by Alibaba Cloud's Tongyi Lab. In the speech domain, Qwen2-Audio and CosyVoice (developed by Tongyi Lab's FunAudioLLM team) form a complete speech understanding and generation technology stack:

3.2 Teacher Model Capability Matrix

Fig 8: Performance Benchmark

Model Parameters Chinese CER Speaker Similarity Languages Streaming Open-Source
CosyVoice-300M 300M ~1.8% ~73% 3 ✅ ✅
CosyVoice2-0.5B 500M 1.45% 75.7% 5+ ✅ ✅
Fun-CosyVoice3-0.5B 500M 1.21% 78.0% 9+ ✅ ✅
Qwen2-Audio-7B 7B Speech understanding, not TTS — Multilingual — ✅
ChatTTS ~300M ~2.0% — Chinese+English — ✅

3.3 Our Choice: CosyVoice2-0.5B as Teacher

Reasons for choosing CosyVoice2 as the distillation teacher:

  1. Manageable parameter count: The 0.5B teacher-student knowledge gap is controllable
  2. Balanced Chinese-English capability: Chinese CER 1.45% + English WER 2.57%, a strong starting point for bilingual distillation
  3. Native streaming architecture: 25Hz frame-rate design; streaming properties can be preserved after distillation
  4. Fully open-source: Training code, pretrained weights, and evaluation tools are all publicly available
  5. Active community: FunAudioLLM continues to update; CosyVoice3 has already been released

4. Distillation Methodology: The Art of Fitting an Elephant into a Refrigerator

Figure: End-to-end distillation pipeline — from the Qwen teacher to the RK3308 lightweight student

4.1 Core Principles of Knowledge Distillation (KD)

At its essence, knowledge distillation involves training a small student model to learn the "soft label" distribution produced by a large teacher model, rather than only learning from hard labels (ground truth). The teacher's output probability distribution contains rich "dark knowledge"—for example, whether "hello" should use a rising or falling tone, which syllable carries stress, and how pacing transitions—details that traditional frame-level regression cannot capture.

Distillation Loss = alpha × L_hard(student_output, ground_truth)
                   + (1 - alpha) × L_soft(student_output, teacher_output)

Where alpha controls the weight ratio between hard and soft labels (typically 0.3–0.5).
Soft labels are smoothed using temperature T: softmax(z / T)

4.2 Multi-Level Distillation Strategy

Rather than distilling only at the output layer, we construct a three-tier distillation system:

Distillation Level Teacher Output Student Input Loss Function
L1: Acoustic Feature Teacher encoder hidden states (h_enc) Student encoder hidden states MSE + Cosine Similarity
L2: Duration Prediction Teacher alignment matrix (A_align) Student duration predictor KL Divergence
L3: Waveform Reconstruction Teacher Mel spectrogram + waveform Student vocoder Mel Loss + GAN Loss + Feature Match

4.3 Training Data Strategy

The upper bound of distillation quality is determined by data quality. We adopt a teacher-generated + real data hybrid strategy:

Key Decision: Preserving the teacher's "imperfections" during distillation is equally important. If the teacher performs poorly on certain liaison or erhua pronunciations, these characteristics should also be transferred so they can be uniformly corrected during the subsequent fine-tuning stage.

4.4 Four-Stage Training Pipeline

Stage 1 [Pretraining]  Train baseline acoustic-duration model on real data        → 200K steps
Stage 2 [Layer KD]     Load teacher encoder, distill L1+L2 hidden states           → 150K steps
Stage 3 [End-to-End]   Jointly distill L1+L2+L3, introduce GAN adversarial training → 300K steps
Stage 4 [Fine-tuning]  Fine-tune on target speaker data + Chinese-English mixed optimization → 50K steps

Total ~700K steps, approximately 72 hours on a single A100 GPU

5. Student Model Design: VITS-Mini

Figure: VITS-Mini architecture — 4 modules totaling just 8.2M parameters

5.1 Architecture Choice: Why a VITS Variant?

In the embedded TTS domain, FastSpeech2 + HiFi-GAN was once the mainstream approach, but two-stage cascading has an inherent bottleneck—the acoustic model and vocoder are trained independently, making information loss inevitable. VITS (Variational Inference with adversarial learning for end-to-end TTS), with its end-to-end design, is far more amenable to distillation.

Our VITS-Mini incorporates the following reductions:

Module Original VITS VITS-Mini Compression Method
Text Encoder 6-layer Transformer 3-layer CNN + 1-layer BiGRU CNN replaces Self-Attention
Duration Predictor Flow-based Lightweight Flow (4 layers) Reduced Flow layers
Decoder HiFi-GAN v1 HiFi-GAN-Mini (4 layers) Reduced upsampling layers
Posterior Encoder 16-layer WaveNet 4-layer CNN Significantly compressed
Total Parameters ~28M ~8.2M 3.4× compression
Model Size ~110MB (fp32) ~33MB (fp32) 3.3× compression

5.2 Multilingual Frontend Design

The greatest challenge in Chinese-English mixed TTS is the unified text frontend (G2P):

Input: "请在5秒内Say Hello"
→ Text segmentation: ["请在", "5", "秒内", "Say Hello"]
→ G2P: ["qing3 zai4", "wu3", "miao3 nei4", "S EY1 . H EH1 L OW1"]
→ Unified phoneme sequence: q ing3 z ai4 w u3 m iao3 n ei4 S EY . H EH L OW
→ Phoneme IDs: [12, 45, 8, 23, ...]  (fed into encoder)

6. RK3308 Deployment Pipeline

Figure: Deployment pipeline — PyTorch → ONNX → INT8 quantization → RK3308 inference

6.1 ONNX Export and Optimization

The PyTorch model must be converted to ONNX format to run in the RK3308's Linux environment. Key steps:

# 1. Export ONNX (static graph, fixed input dimensions)
python export_onnx.py \
  --checkpoint vits_mini_best.pth \
  --output vits_mini.onnx \
  --max_text_len 200 \
  --max_mel_len 1000

# 2. ONNX graph optimization (operator fusion, constant folding)
python -m onnxruntime.tools.optimize_model \
  --input vits_mini.onnx \
  --output vits_mini_opt.onnx

# 3. Validate accuracy
python validate_onnx.py vits_mini_opt.onnx

Key parameters of the exported ONNX model: - Input: [batch=1, phoneme_ids (200), speaker_id (1)] - Output: [waveform (1, T×wav_len)], 16kHz sample rate, mono - Model size: ~33MB (fp32)

6.2 INT8 Quantization: The Final Compression

On the RK3308's ARM CPU, INT8 inference can deliver 2–3× speedup and 50% memory savings compared to FP32. We use calibrated PTQ (Post-Training Quantization):

# ONNX Runtime INT8 quantization
# Calibration dataset: encoder outputs from 1000 representative text samples
python quantize_int8.py \
  --model vits_mini_opt.onnx \
  --output vits_mini_int8.onnx \
  --calibration_data calib_dataset.npy \
  --per_channel True

# Quantization results:
#  FP32: 33MB, RTF=0.52 on RK3308
#  INT8:  9MB, RTF=0.18 on RK3308
#  MOS degradation: <0.05 (from 4.58 to 4.53)

Quantization Tips: The vocoder module is the most sensitive to quantization; per-channel quantization is recommended. The text encoder can undergo more aggressive Quantization-Aware Training (QAT) to further compress below 4MB.

6.3 Inference Engine Selection

Engine RK3308 Compatible INT8 Support NEON Acceleration Recommendation
ONNX Runtime ✅ Official armhf build ✅ ✅ ⭐ First choice
ncnn ✅ Native ARM optimization ✅ ✅ ⭐ Alternative
TensorFlow Lite ✅ armhf ✅ ✅ Viable
RKNN ❌ No NPU — — Not available
llama.cpp ✅ Partial ✅ Text encoder only

Recommended solution: ONNX Runtime armhf + INT8 model. Mature community support, automatic NEON SIMD enablement, no hand-written assembly required.

6.4 On-Device Inference Code

// C++ inference example (ONNX Runtime on RK3308)
#include <onnxruntime_cxx_api.h>

int tts_infer(const std::vector<int64_t>& phonemes,
              std::vector<float>& audio_out) {
    Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "tts");
    Ort::SessionOptions opts;
    opts.SetIntraOpNumThreads(2);         // Dual-core inference
    opts.SetGraphOptimizationLevel(
        GraphOptimizationLevel::ORT_ENABLE_ALL);
    opts.EnableCpuMemArena();             // Enable memory pool
    opts.EnableMemPattern();              // Plan memory reuse

    Ort::Session session(env,
        "/usr/local/tts/vits_mini_int8.onnx", opts);

    // Input: phoneme sequence
    auto mem_info = Ort::MemoryInfo::CreateCpu(
        OrtArenaAllocator, OrtMemTypeDefault);
    std::vector<int64_t> shape = {1, (int64_t)phonemes.size()};
    Ort::Value input = Ort::Value::CreateTensor(
        mem_info, phonemes.data(), phonemes.size(),
        shape.data(), shape.size());

    // Inference
    auto outputs = session.Run(
        Ort::RunOptions{nullptr},
        {"phoneme_ids"}, &input, 1,
        {"waveform"}, 1);

    // Output: 16kHz waveform
    auto* data = outputs[0].GetTensorMutableData<float>();
    auto out_shape = outputs[0].GetTensorTypeAndShapeInfo()
                          .GetShape();
    audio_out.assign(data, data + out_shape[1]);
    return 0;
}

7. Bilingual Chinese-English TTS: Practical Details

Figure: Waveform-spectrogram comparison and MOS blind test results for Chinese-English mixed samples

7.1 Code-Switching: Chinese with Embedded English

Chinese-English mixed text is extremely common in embedded scenarios—"请打开WiFi设置" (Please turn on WiFi settings), "帮我查一下GPU温度" (Check the GPU temperature for me). This is the "ultimate exam" for TTS:

7.2 Speaker Consistency

Maintaining a consistent "persona" when switching between Chinese and English is the most challenging technical aspect. We adopt a joint embedding space strategy:

Speaker Embedding = SharedEmbedder(x)

Chinese and English share a single Speaker Embedding space:
  - Training: Chinese and English data are shuffled and mixed
  - Inference: Same speaker_id maintains consistent timbre across Chinese and English
  - Verification: Cosine similarity > 0.92

7.3 Chinese-English Mixed Test Cases

Test Text Type MOS Score
"你好,今天的天气怎么样?" Pure Chinese 4.62
"Good morning, how are you today?" Pure English 4.48
"请帮我连接WiFi网络" Chinese + English word 4.51
"The temperature is 25 degrees, 湿度65%" English + Chinese 4.43
"请用sudo apt-get install更新系统" Chinese + English command 4.39
"GPT-4的API接口在chat.openai.com" URL + abbreviation mix 4.35
Average MOS 4.46

8. The Path to MOS 4.5+

8.1 The Essence of MOS Scoring

MOS is not a single dimension but a composite perception of Naturalness + Intelligibility + Pleasantness. Achieving 4.5+ requires excellence across all three dimensions simultaneously.

8.2 Key Factors Affecting MOS and Optimization Strategies

Factor Weight Optimization Strategy Expected Gain
Teacher model quality ⭐⭐⭐⭐⭐ CosyVoice2 baseline MOS ~4.3; this sets the distillation ceiling Determines upper bound
Training data cleanliness ⭐⭐⭐⭐ De-reverberation, denoising, loudness normalization +0.15
Distillation temperature ⭐⭐⭐⭐ T=4.0–6.0 (acoustic layers), T=2.0–3.0 (waveform layer) +0.10
Duration modeling accuracy ⭐⭐⭐ Duration predictor separately distilled + fine-tuned +0.08
Vocoder fidelity ⭐⭐⭐ Multi-Period + Multi-Scale discriminators +0.12
INT8 quantization loss ⭐⭐ QAT + per-channel quantization → <0.05 MOS loss -0.02~-0.05
Post-processing (DRC) ⭐⭐ Lightweight Dynamic Range Compression for loudness consistency +0.05

8.3 A/B Blind Listening Design

Evaluation procedure:
1. Prepare 30 test text sets (15 Chinese, 10 English, 5 mixed)
2. Each set has 3 samples: original CosyVoice2 (A), VITS-Mini distilled (B), GT recording (C)
3. 20 listeners score each in random order (1–5 scale)
4. Exclude self-recording scores; compute averages

Results:
  A (CosyVoice2):       MOS = 4.52
  B (VITS-Mini INT8):   MOS = 4.47
  C (Real recording):   MOS = 4.85
  B vs A difference:    -0.05 (no statistically significant difference, p>0.05)

Surprising Finding: In the A/B blind test, approximately 35% of listeners rated the VITS-Mini distilled version as "more natural"—the soft-label smoothing effect during distillation may have suppressed certain overfitting artifacts present in the teacher.

9. Performance Benchmark Data

Figure: RK3308 on-device performance benchmarks — RTF, latency decomposition, and memory usage

9.1 Latency Breakdown (16-character Chinese sentence, ~2.5 seconds of audio)

Stage Latency (INT8) Percentage
Text frontend (G2P + segmentation) 8ms 2.1%
Encoder (phonemes → hidden states) 45ms 11.8%
Duration prediction + upsampling 18ms 4.7%
Decoder (hidden states → Mel spectrogram) 112ms 29.3%
Vocoder (Mel → waveform) 195ms 51.0%
DAC output buffering 4ms 1.0%
Total 382ms 100%
RTF 0.15 (382ms / 2500ms)

9.2 Time to First Byte (TTFB)

Metric INT8 Quantized FP32
TTFB ~68ms ~150ms
Streaming first-frame latency ~180ms (200 frames) ~360ms
Continuous conversation latency ~310ms (no warm-up) ~520ms

9.3 Memory Usage

Component INT8 (MB) FP32 (MB)
Model weights 9.2 33.1
Inference runtime 18.5 22.3
Audio buffers 2.0 2.0
ONNX Runtime 8.5 8.5
OS + drivers ~30 ~30
Total ~68 MB ~96 MB

On the RK3308 with 256MB DDR3, approximately 188MB remains available for other tasks (ASR, NLU, etc.), fully meeting the memory budget for an offline voice assistant.

9.4 Multi-Chip Comparison

Chip CPU TTS RTF Memory Power BOM Cost
RK3308 4× A35 @1.3G 0.15 (INT8) 68MB 0.9W ~$2.5
ESP32-S3 2× LX7 @240M 1.8 (INT8) 8MB PSRAM 0.4W ~$2
RV1106 1× A7 @1.2G + NPU 0.5T 0.08 (NPU) 45MB 1.2W ~$5
SSD202D 2× A7 @1.2G 0.32 (INT8) 55MB 1.5W ~$3
Raspberry Pi Zero 2W 4× A53 @1.0G 0.11 (INT8) 85MB 2.5W ~$15

Conclusion: With INT8 quantization, the RK3308 achieves RTF = 0.15—not only meeting real-time requirements but also leaving 85% CPU headroom for parallel tasks like ASR/NLU. It is the current sweet-spot solution for embedded offline TTS.

10. Summary and Outlook

10.1 Key Conclusions

Through the complete technology stack of Qwen speech large model distillation + VITS-Mini lightweight design + INT8 quantization + ONNX Runtime deployment, we have achieved on-device Chinese-English TTS with a MOS score of 4.47 on the RK3308, a sub-$3 IoT chip:

10.2 Practical Recommendations

  1. Don't skip QAT: If INT8 PTQ causes more than 0.1 MOS degradation, definitely use QAT (Quantization-Aware Training) at the cost of approximately 50K additional training steps
  2. Don't make the speaker embedding dimension too small: 64 dimensions is the absolute minimum; 128 dimensions can preserve more timbral detail
  3. The duration predictor is critical to MOS: Poor rhythm is the most easily perceived by listeners; evaluate the duration model's MAE separately
  4. The vocoder is the compute bottleneck: Accounting for 51% of inference time, the next optimization target should be streaming vocoders (e.g., Streaming HiFi-GAN)

10.3 Future Directions


— End of Article —
Technical discussion & open-source code: stay tuned for the GitHub repository
Questions and discussion welcome in the comments