✨ Qwen3-ASR-1.7B-334h

A LoRA adapter for Qwen/Qwen3-ASR-1.7B-hf, fine-tuned on 334 hours of Vietnamese speech drawn in roughly equal parts from four corpora β€” viVoice (YouTube), VietSpeech (social media), VieNeu-TTS (studio read speech) and Bud500 (YouTube, short chunks).

69 MB of adapter weights: 17.4 M trainable parameters (0.85% of the base), on the 28 decoder layers only β€” all four attention projections plus all three MLP projections, 196 modules. The audio encoder is frozen.

Every corpus improves, and every test split is genuinely held out. Each corpus keeps its own speaker- or recording-disjoint split, and training used only the train sides. WER falls 13–62% relative, and all four deltas clear a bootstrap 95% CI. That is the difference between this adapter and its 34 h sibling, whose external benchmarks were in-domain after fine-tuning.

Trained on short clips (mean 2.5–6.6 s, longest 16.1 s), but tested on long ones: on 5–60 s segments it does not regress, and it does not truncate. See 🎯 Accuracy β†’ Long-form.

⚠️ Punctuation is domain-conditional β€” the adapter reproduces whichever convention its training corpus used for that kind of audio. See πŸ–‹ Output style.


πŸš‹ Usage

Below is a quick way to get up and running with the model.

  1. Install
pip install "transformers>=5.14" peft torch torchaudio soundfile librosa

transformers>=5.14 is required β€” the qwen3_asr architecture was added after 5.5.x, and earlier versions raise KeyError: 'qwen3_asr'.

  1. Load the base model and attach the adapter
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
from peft import PeftModel

BASE = "Qwen/Qwen3-ASR-1.7B-hf"
ADAPTER = "tyanfarm/Qwen3-ASR-1.7B-334h"

processor = AutoProcessor.from_pretrained(BASE)
model = AutoModelForMultimodalLM.from_pretrained(
    BASE, dtype=torch.bfloat16, attn_implementation="sdpa", device_map="cuda",
)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()

# Batched generation is only correct with left padding.
processor.tokenizer.padding_side = "left"
  1. Transcribe
import librosa

# Pass a decoded 16 kHz float array, not a file path: the processor's own file
# loader goes through torchcodec+FFmpeg, and FFmpeg 4 hits a "0 channels" bug.
audio, _ = librosa.load("sample.wav", sr=16000)

inputs = processor.apply_transcription_request(
    audio=[audio], language="Vietnamese",
).to(model.device, model.dtype)

with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=440, do_sample=False)

gen = out[:, inputs["input_ids"].shape[1]:]
print(processor.decode(gen, return_format="transcription_only")[0])

Long audio: chunk to 60 s or less, and use a VAD so chunks land on silence rather than mid-word. 5–60 s is tested (see 🎯 Accuracy β†’ Long-form); past ~75 s the default max_new_tokens=440 silently truncates the hypothesis.

Batching. With padding_side="left" set above, audio=[a1, a2, ...] transcribes several clips per call. Sort by duration first β€” mixing a 16 s clip with a 2 s one spends most of the batch on padding. On 16 GB, use batch 1–2 once clips approach 60 s (measured 6.7 GB peak on a single 57 s clip).


πŸ“‘ Training data

Four corpora, each capped at 100 h independently so no single source dominates the gradient. The caps were not all reached; these are the measured cached hours, 315,029 clips in the train split:

corpus source character clips hours
bud500 linhtran92/viet_bud500 YouTube, fixed-length chunks cut mid-phrase 126,326 89.95
vietspeech NhutP/VietSpeech social media, 3 regional accents 74,014 82.53
vieneu pnnbao-ump/VieNeu-TTS-140h studio TTS read speech 43,451 81.61
vivoice_full capleaf/viVoice YouTube, native unmerged clips 71,238 80.31
total 315,029 334.39

Splits are speaker- or recording-disjoint, and built per corpus. None of the four ships an official test split, so they were constructed β€” never at the clip level, which would put neighbouring clips from one recording on both sides and report a WER the model did not earn. viVoice and Bud500 split by channel, VietSpeech by the recording prefix in its filenames, VieNeu by voice (its speaker column holds one id per recording, so the trailing index is stripped first). Training used only the train sides; the 🎯 numbers below are the test sides.

Hyperparameters

LoRA rank / alpha / dropout 16 / 32 / 0.05
target modules language_model.layers.*.{q,k,v,o}_proj + mlp.{gate,up,down}_proj (196)
trainable params 17,432,576 (0.85% of 2.06 B)
epochs / optimizer steps 1 / 19,690
batch 8 Γ— 2 grad-accum (effective 16), group_by_length
learning rate 1e-4, cosine, 3% warmup
precision bf16, gradient checkpointing, adamw_torch_fused
checkpoint selection load_best_model_at_end on eval_loss
hardware 1 Γ— RTX 5080 16 GB, 11.73 GB peak VRAM

Validation loss fell monotonically across all ten evals β€” 0.1415 (step 2,000) β†’ 0.1117 (step 19,690) β€” so the best checkpoint is the last one, and the single epoch ended still improving. A second epoch was not run.


🎯 Accuracy

WER/CER are computed after Vietnamese normalization that expands digits to their spoken form on both sides (334% → ba trăm ba mưƑi bốn phần trăm) and folds the valid alternative readings (tư/bốn, mốt/một, lăm/năm, ngàn/nghìn, linh/lẻ). Without this, one formatting mismatch can cost 100% WER on a short utterance.

Each corpus's own held-out test split, 500 clips sampled per corpus, scored against the stock base model on the identical clips:

corpus n base WER LoRA WER abs rel 95% CI on the delta
bud500 500 4.98% 1.90% βˆ’3.08 βˆ’61.8% [βˆ’3.75, βˆ’2.45] βœ”
vieneu 500 3.92% 3.08% βˆ’0.83 βˆ’21.3% [βˆ’1.27, βˆ’0.35] βœ”
vivoice_full 500 5.43% 4.67% βˆ’0.76 βˆ’14.0% [βˆ’1.13, βˆ’0.41] βœ”
vietspeech 500 6.59% 5.73% βˆ’0.86 βˆ’13.1% [βˆ’1.32, βˆ’0.43] βœ”

CER moves the same way: 3.62% β†’ 0.96% (bud500), 1.97% β†’ 1.66% (vieneu), 2.70% β†’ 2.27% (vivoice_full), 3.76% β†’ 3.44% (vietspeech).

All four intervals exclude zero β€” bootstrap over the per-clip errors, so these are not noise. Compare with this project's 34 h run, where a 0.02-point "win" on viVoice turned out to be 2 word errors in 10,155.

Why Bud500 moves 4Γ— further than the rest. Its clips are fixed-length YouTube chunks cut mid-phrase (mean 2.5 s, max 4.5 s) with bare lowercase transcripts β€” the format furthest from what the base model was pretrained to emit, so it has the most format mismatch to recover. Read βˆ’61.8% as "the adapter learned this corpus's conventions", not as a claim about Vietnamese ASR difficulty.

Long-form: training on short clips did not break it

Every number above scores clips under 16 s, so none of them can see whether a LoRA trained on 2–6 s utterances learned "speech ends after a few seconds" and started truncating on long input. This is the test that can: 114 merged viVoice segments of 5–60 s (mean 21.5 s), same normalizer, scored against the stock model on the identical clips.

slice n base WER LoRA WER delta 95% CI
unseen speakers, overall 45 5.97% 5.52% βˆ’0.45 [βˆ’1.08, +0.17] β€” noise
unseen, 5–30 s 35 5.95% 5.51% βˆ’0.44 [βˆ’1.28, +0.40] β€” noise
unseen, 30–60 s 10 6.01% 5.55% βˆ’0.47 [βˆ’1.44, +0.52] β€” noise
in-domain, overall 114 5.96% 5.13% βˆ’0.83 [βˆ’1.26, βˆ’0.41] βœ”
in-domain, 30–60 s 20 6.09% 5.24% βˆ’0.85 [βˆ’1.64, βˆ’0.14] βœ”

Read the top three rows as the result: no regression on long audio. 13 of the 17 channels in this split are inside vivoice_full's training data β€” both caches stream viVoice from the head of the same split β€” so the 114-clip rows are an in-domain probe, and their gain is concentrated in exactly those channels (seen-only: βˆ’1.09 pts). The 45 clips on the 4 unseen channels are the honest slice, and there the interval spans zero: the adapter neither helps nor hurts long-form for new speakers.

No truncation, no looping. LoRA hypotheses average 1.001Γ— the reference word count on 30–60 s clips (base: 1.015Γ—), no clip fell below 0.92Γ— on any bucket, and the longest segment (57.9 s) returned 231 words against a 231-word reference, ending on a complete sentence. That was the failure mode this test existed to catch, and it does not appear.


πŸ–‹ Output style

The adapter reproduces whichever transcript convention its training corpus used for that kind of audio. Share of hypotheses containing any of . , ? !:

test set source refs base model this adapter
vivoice_full 100% punctuated 93.8% 100.0%
vieneu 100% punctuated 97.4% 100.0%
vietspeech 0% punctuated 92.8% 0.0%
bud500 0% punctuated 77.4% 0.0%

The mapping from source convention to output is exact, in both directions. Leading capitalization follows it: 42.8% β†’ 100.0% on viVoice and 27.6% β†’ 100.0% on VieNeu, against 17.8% β†’ 12.4% on VietSpeech and 10.0% β†’ 7.8% on Bud500.

So the adapter did not learn "punctuate" or "don't punctuate" β€” it learned to predict the convention from the acoustic domain, because the four corpora disagree and the audio tells them apart. Studio and YouTube-narration audio comes back fully punctuated and capitalized; conversational and short-chunk audio comes back bare.

This does not affect the WER/CER tables above β€” the metric strips punctuation and lowercases before scoring. It matters only if you consume the transcript directly. If you need punctuation guaranteed on conversational audio, run a punctuation restoration model over the output.


⚠️ Limitations

  • Long-form is verified to 60 s, not beyond. The 5–60 s test above shows no regression and no truncation, but nothing here scores 2–10 minute audio; at that tier max_new_tokens=440 becomes the binding limit (it covers ~75 s at the measured 5.8 tokens/s p95) and you would chunk anyway. Raise it in step with duration if you go longer.
  • The long-form gain is in-domain. On genuinely unseen speakers the long-form delta is not distinguishable from noise at n=45. "No regression" is the claim that slice supports; "improves long-form" is not.
  • Vietnamese only. Trained with language="Vietnamese" on every request; other languages are untested and likely degraded.
  • Punctuation is domain-conditional β€” see πŸ–‹ Output style above.
  • Bud500 is 40% of the training clips but only 27% of the hours, because its clips are the shortest. Clip-count-weighted, the mixture leans further toward very short utterances than the hours table suggests.
  • r=16 on 315 k examples is on the small side. The single epoch ended with validation loss still falling; r=32/r=64 or a second epoch is the obvious next thing to try.
  • Scored at 500 clips per corpus, not the full test splits.

πŸ“„ License

cc-by-nc-sa-4.0. The base model is Apache-2.0 and two of the four corpora (VietSpeech, VieNeu-TTS) are Apache-2.0 β€” but viVoice and Bud500 are both CC BY-NC-SA 4.0, and this adapter is a derivative of all four, so the non-commercial share-alike terms carry over. Use it for research and personal projects; commercial use would require re-training on the permissive subset.

Downloads last month
23
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for tyanfarm/Qwen3-ASR-1.7B-334h

Adapter
(12)
this model

Datasets used to train tyanfarm/Qwen3-ASR-1.7B-334h