Instructions to use tyanfarm/Qwen3-ASR-1.7B-334h with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use tyanfarm/Qwen3-ASR-1.7B-334h with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-ASR-1.7B-hf") model = PeftModel.from_pretrained(base_model, "tyanfarm/Qwen3-ASR-1.7B-334h") - Notebooks
- Google Colab
- Kaggle
β¨ Qwen3-ASR-1.7B-334h
A LoRA adapter for Qwen/Qwen3-ASR-1.7B-hf,
fine-tuned on 334 hours of Vietnamese speech drawn in roughly equal parts from
four corpora β viVoice (YouTube), VietSpeech (social media), VieNeu-TTS (studio
read speech) and Bud500 (YouTube, short chunks).
69 MB of adapter weights: 17.4 M trainable parameters (0.85% of the base), on the 28 decoder layers only β all four attention projections plus all three MLP projections, 196 modules. The audio encoder is frozen.
Every corpus improves, and every test split is genuinely held out. Each corpus keeps its own speaker- or recording-disjoint split, and training used only the train sides. WER falls 13β62% relative, and all four deltas clear a bootstrap 95% CI. That is the difference between this adapter and its 34 h sibling, whose external benchmarks were in-domain after fine-tuning.
Trained on short clips (mean 2.5β6.6 s, longest 16.1 s), but tested on long ones: on 5β60 s segments it does not regress, and it does not truncate. See π― Accuracy β Long-form.
β οΈ Punctuation is domain-conditional β the adapter reproduces whichever convention its training corpus used for that kind of audio. See π Output style.
π Usage
Below is a quick way to get up and running with the model.
- Install
pip install "transformers>=5.14" peft torch torchaudio soundfile librosa
transformers>=5.14 is required β the qwen3_asr architecture was added after
5.5.x, and earlier versions raise KeyError: 'qwen3_asr'.
- Load the base model and attach the adapter
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
from peft import PeftModel
BASE = "Qwen/Qwen3-ASR-1.7B-hf"
ADAPTER = "tyanfarm/Qwen3-ASR-1.7B-334h"
processor = AutoProcessor.from_pretrained(BASE)
model = AutoModelForMultimodalLM.from_pretrained(
BASE, dtype=torch.bfloat16, attn_implementation="sdpa", device_map="cuda",
)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()
# Batched generation is only correct with left padding.
processor.tokenizer.padding_side = "left"
- Transcribe
import librosa
# Pass a decoded 16 kHz float array, not a file path: the processor's own file
# loader goes through torchcodec+FFmpeg, and FFmpeg 4 hits a "0 channels" bug.
audio, _ = librosa.load("sample.wav", sr=16000)
inputs = processor.apply_transcription_request(
audio=[audio], language="Vietnamese",
).to(model.device, model.dtype)
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=440, do_sample=False)
gen = out[:, inputs["input_ids"].shape[1]:]
print(processor.decode(gen, return_format="transcription_only")[0])
Long audio: chunk to 60 s or less, and use a VAD so chunks land on silence
rather than mid-word. 5β60 s is tested (see π― Accuracy β Long-form); past ~75 s
the default max_new_tokens=440 silently truncates the hypothesis.
Batching. With padding_side="left" set above, audio=[a1, a2, ...]
transcribes several clips per call. Sort by duration first β mixing a 16 s clip
with a 2 s one spends most of the batch on padding. On 16 GB, use batch 1β2 once
clips approach 60 s (measured 6.7 GB peak on a single 57 s clip).
π‘ Training data
Four corpora, each capped at 100 h independently so no single source dominates the gradient. The caps were not all reached; these are the measured cached hours, 315,029 clips in the train split:
| corpus | source | character | clips | hours |
|---|---|---|---|---|
bud500 |
linhtran92/viet_bud500 |
YouTube, fixed-length chunks cut mid-phrase | 126,326 | 89.95 |
vietspeech |
NhutP/VietSpeech |
social media, 3 regional accents | 74,014 | 82.53 |
vieneu |
pnnbao-ump/VieNeu-TTS-140h |
studio TTS read speech | 43,451 | 81.61 |
vivoice_full |
capleaf/viVoice |
YouTube, native unmerged clips | 71,238 | 80.31 |
| total | 315,029 | 334.39 |
Splits are speaker- or recording-disjoint, and built per corpus. None of the
four ships an official test split, so they were constructed β never at the clip
level, which would put neighbouring clips from one recording on both sides and
report a WER the model did not earn. viVoice and Bud500 split by channel,
VietSpeech by the recording prefix in its filenames, VieNeu by voice (its
speaker column holds one id per recording, so the trailing index is stripped
first). Training used only the train sides; the π― numbers below are the test
sides.
Hyperparameters
| LoRA rank / alpha / dropout | 16 / 32 / 0.05 |
| target modules | language_model.layers.*.{q,k,v,o}_proj + mlp.{gate,up,down}_proj (196) |
| trainable params | 17,432,576 (0.85% of 2.06 B) |
| epochs / optimizer steps | 1 / 19,690 |
| batch | 8 Γ 2 grad-accum (effective 16), group_by_length |
| learning rate | 1e-4, cosine, 3% warmup |
| precision | bf16, gradient checkpointing, adamw_torch_fused |
| checkpoint selection | load_best_model_at_end on eval_loss |
| hardware | 1 Γ RTX 5080 16 GB, 11.73 GB peak VRAM |
Validation loss fell monotonically across all ten evals β 0.1415 (step 2,000) β 0.1117 (step 19,690) β so the best checkpoint is the last one, and the single epoch ended still improving. A second epoch was not run.
π― Accuracy
WER/CER are computed after Vietnamese normalization that expands digits to their
spoken form on both sides (334% β ba trΔm ba mΖ°Ζ‘i bα»n phαΊ§n trΔm) and folds
the valid alternative readings (tΖ°/bα»n, mα»t/mα»t, lΔm/nΔm,
ngà n/nghìn, linh/lẻ). Without this, one formatting mismatch can cost 100%
WER on a short utterance.
Each corpus's own held-out test split, 500 clips sampled per corpus, scored against the stock base model on the identical clips:
| corpus | n | base WER | LoRA WER | abs | rel | 95% CI on the delta |
|---|---|---|---|---|---|---|
bud500 |
500 | 4.98% | 1.90% | β3.08 | β61.8% | [β3.75, β2.45] β |
vieneu |
500 | 3.92% | 3.08% | β0.83 | β21.3% | [β1.27, β0.35] β |
vivoice_full |
500 | 5.43% | 4.67% | β0.76 | β14.0% | [β1.13, β0.41] β |
vietspeech |
500 | 6.59% | 5.73% | β0.86 | β13.1% | [β1.32, β0.43] β |
CER moves the same way: 3.62% β 0.96% (bud500), 1.97% β 1.66% (vieneu), 2.70% β 2.27% (vivoice_full), 3.76% β 3.44% (vietspeech).
All four intervals exclude zero β bootstrap over the per-clip errors, so these are not noise. Compare with this project's 34 h run, where a 0.02-point "win" on viVoice turned out to be 2 word errors in 10,155.
Why Bud500 moves 4Γ further than the rest. Its clips are fixed-length YouTube chunks cut mid-phrase (mean 2.5 s, max 4.5 s) with bare lowercase transcripts β the format furthest from what the base model was pretrained to emit, so it has the most format mismatch to recover. Read β61.8% as "the adapter learned this corpus's conventions", not as a claim about Vietnamese ASR difficulty.
Long-form: training on short clips did not break it
Every number above scores clips under 16 s, so none of them can see whether a LoRA trained on 2β6 s utterances learned "speech ends after a few seconds" and started truncating on long input. This is the test that can: 114 merged viVoice segments of 5β60 s (mean 21.5 s), same normalizer, scored against the stock model on the identical clips.
| slice | n | base WER | LoRA WER | delta | 95% CI |
|---|---|---|---|---|---|
| unseen speakers, overall | 45 | 5.97% | 5.52% | β0.45 | [β1.08, +0.17] β noise |
| unseen, 5β30 s | 35 | 5.95% | 5.51% | β0.44 | [β1.28, +0.40] β noise |
| unseen, 30β60 s | 10 | 6.01% | 5.55% | β0.47 | [β1.44, +0.52] β noise |
| in-domain, overall | 114 | 5.96% | 5.13% | β0.83 | [β1.26, β0.41] β |
| in-domain, 30β60 s | 20 | 6.09% | 5.24% | β0.85 | [β1.64, β0.14] β |
Read the top three rows as the result: no regression on long audio. 13 of the
17 channels in this split are inside vivoice_full's training data β both caches
stream viVoice from the head of the same split β so the 114-clip rows are an
in-domain probe, and their gain is concentrated in exactly those channels
(seen-only: β1.09 pts). The 45 clips on the 4 unseen channels are the honest
slice, and there the interval spans zero: the adapter neither helps nor hurts
long-form for new speakers.
No truncation, no looping. LoRA hypotheses average 1.001Γ the reference word count on 30β60 s clips (base: 1.015Γ), no clip fell below 0.92Γ on any bucket, and the longest segment (57.9 s) returned 231 words against a 231-word reference, ending on a complete sentence. That was the failure mode this test existed to catch, and it does not appear.
π Output style
The adapter reproduces whichever transcript convention its training corpus used
for that kind of audio. Share of hypotheses containing any of . , ? !:
| test set | source refs | base model | this adapter |
|---|---|---|---|
vivoice_full |
100% punctuated | 93.8% | 100.0% |
vieneu |
100% punctuated | 97.4% | 100.0% |
vietspeech |
0% punctuated | 92.8% | 0.0% |
bud500 |
0% punctuated | 77.4% | 0.0% |
The mapping from source convention to output is exact, in both directions. Leading capitalization follows it: 42.8% β 100.0% on viVoice and 27.6% β 100.0% on VieNeu, against 17.8% β 12.4% on VietSpeech and 10.0% β 7.8% on Bud500.
So the adapter did not learn "punctuate" or "don't punctuate" β it learned to predict the convention from the acoustic domain, because the four corpora disagree and the audio tells them apart. Studio and YouTube-narration audio comes back fully punctuated and capitalized; conversational and short-chunk audio comes back bare.
This does not affect the WER/CER tables above β the metric strips punctuation and lowercases before scoring. It matters only if you consume the transcript directly. If you need punctuation guaranteed on conversational audio, run a punctuation restoration model over the output.
β οΈ Limitations
- Long-form is verified to 60 s, not beyond. The 5β60 s test above shows no
regression and no truncation, but nothing here scores 2β10 minute audio; at that
tier
max_new_tokens=440becomes the binding limit (it covers ~75 s at the measured 5.8 tokens/s p95) and you would chunk anyway. Raise it in step with duration if you go longer. - The long-form gain is in-domain. On genuinely unseen speakers the long-form delta is not distinguishable from noise at n=45. "No regression" is the claim that slice supports; "improves long-form" is not.
- Vietnamese only. Trained with
language="Vietnamese"on every request; other languages are untested and likely degraded. - Punctuation is domain-conditional β see π Output style above.
- Bud500 is 40% of the training clips but only 27% of the hours, because its clips are the shortest. Clip-count-weighted, the mixture leans further toward very short utterances than the hours table suggests.
r=16on 315 k examples is on the small side. The single epoch ended with validation loss still falling;r=32/r=64or a second epoch is the obvious next thing to try.- Scored at 500 clips per corpus, not the full test splits.
π License
cc-by-nc-sa-4.0. The base model is Apache-2.0 and two of the four corpora
(VietSpeech, VieNeu-TTS) are Apache-2.0 β but viVoice and Bud500 are both
CC BY-NC-SA 4.0, and this adapter is a derivative of all four, so the
non-commercial share-alike terms carry over. Use it for research and personal
projects; commercial use would require re-training on the permissive subset.
- Downloads last month
- 23
Model tree for tyanfarm/Qwen3-ASR-1.7B-334h
Base model
Qwen/Qwen3-ASR-1.7B-hf