๐ Quick Start โข ๐๏ธ 21 Voice Roster โข ๐ Multilingual Routing โข ๐ Architecture โข โก Benchmarks
๐ Why Resona?
Traditional TTS engines either require expensive cloud APIs with latency spikes or generate robotic, metallic voices when dealing with Indian languages and Hinglish code-switching. Resona solves this with an ultra-compact 82M parameter neural architecture running 100% offline on standard consumer CPUs.
๐ฎ๐ณ Native Bilingual Code-SwitchingSeamlessly transitions between English UI / technical terminology and native Hindi diction without unnatural accent distortion or robotic stutter. |
โก Real-Time Edge LatencyGenerates speech in sub-second time on standard commodity CPUs (RTF < 0.20x). Zero GPU or specialized hardware required. |
๐ซ Natural Breathing CadenceBuilt-in biological rhythm injector that inserts subtle 180ms clause pauses and 350ms sentence boundaries for conversational realism. |
๐ 100% Private & Air-GappedZero external API calls, zero telemetry, and zero data leakage. All weights, style vectors, and lexicons run locally on your device. |
๐ Pipeline Architecture
flowchart LR
A[Raw Input Text] --> B[Text Normalizer]
B --> C{Language Router}
C -- Hinglish --> D[Hinglish Transducer]
D --> E[Protected Tech Whitelist Guard]
E --> F[Devanagari / IPA Phonemizer]
C -- Multilingual / English --> F
F --> G[Token & Rhythm Alignment]
H[Voice Style Tensor 256-D] --> I[AdaIN Prosody Predictor]
G --> I
I --> J[82M Resona Neural Vocoder]
J --> K[Natural Breathing Cadence]
K --> L[Studio WAV Audio 24kHz]
style A fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
style D fill:#313244,stroke:#f38ba8,stroke-width:2px,color:#cdd6f4
style E fill:#313244,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4
style J fill:#45475a,stroke:#cba6f7,stroke-width:3px,color:#f5e0dc
style L fill:#1e1e2e,stroke:#a6e3a1,stroke-width:2px,color:#a6e3a1
๐ Quick Start
1. Installation
# Clone the repository
git clone https://github.com/MilannSharma/Resona.git
cd Resona
# Install package dependencies
pip install -e .
Note: Requires
espeak-nginstalled on your host system (standard package on Linux/macOS, or via installer on Windows). Resona includes automatic multi-path discovery and CLI fallback.
2. Python API
from resona import ResonaPipeline
# 1. Initialize pipeline with flagship Hinglish voice (Anjura)
pipeline = ResonaPipeline(voice="anjura", language="hinglish")
# 2. Synthesize with natural code-switching & breathing cadence
text = "Resona audio engine se speech generate karna super fast aur natural hai. File menu par click kijiye."
result = pipeline.synthesize(text, output_path="output.wav", speed=0.95)
print(f"Generated {result.duration_seconds:.2f}s studio audio at {result.sample_rate}Hz!")
3. Terminal CLI (resona-tts)
# List all 21 available Grade A personas
resona-tts --list-voices
# Synthesize speech directly from command line
resona-tts --text "Welcome to the Resona Neural Audio Engine." \
--voice arjun \
--output welcome.wav
# Synthesize Hinglish tutorial with custom speed multiplier
resona-tts --text "Settings menu par click kijiye aur resolution 1080p select kijiye." \
--voice anjura \
--lang hinglish \
--speed 0.95 \
--output settings.wav
๐๏ธ Studio Voice Roster (21 Personas)
All 21 voices include pre-rendered, high-fidelity reference .wav previews in resona/voices/samples/:
๐ฎ๐ณ 1. Hinglish & Indic Flagships (Studio Curated)
| Voice ID | Display Name | Gender | Persona Character | Pitch | Speed | Preview |
|---|---|---|---|---|---|---|
anjura |
Anjura | ๐ฉ | Flagship Corporate Explainer | 184 Hz | 0.95x | โถ๏ธ Preview |
divya |
Divya | ๐ฉ | Deep, Articulate Explainer | 184.6 Hz | 0.95x | โถ๏ธ Preview |
meera |
Meera | ๐ฉ | Warm Conversational Native Hindi | 210 Hz | 1.00x | โถ๏ธ Preview |
priya |
Priya | ๐ฉ | Dynamic Tech Educator | 195 Hz | 1.00x | โถ๏ธ Preview |
arjun |
Arjun | ๐จ | Enterprise Baritone Narrator | 115 Hz | 0.95x | โถ๏ธ Preview |
kabir |
Kabir | ๐จ | Executive Conversational Podcast | 128 Hz | 0.95x | โถ๏ธ Preview |
aman |
Aman | ๐จ | Agile Tech Walkthrough (Denoised) | 132 Hz | 1.00x | โถ๏ธ Preview |
atul |
Atul | ๐จ | Classic Indic Corporate Voice | 122 Hz | 0.98x | โถ๏ธ Preview |
๐๏ธ 2. Indic English & Native Hindi Series
| Voice ID | Display Name | Gender | Persona Character | Pitch | Speed | Preview |
|---|---|---|---|---|---|---|
ananya |
Ananya | ๐ฉ | Friendly Academic Explainer | 190 Hz | 1.00x | โถ๏ธ Preview |
nisha |
Nisha | ๐ฉ | Calm, Gentle Corporate Presenter | 182 Hz | 0.98x | โถ๏ธ Preview |
tara |
Tara | ๐ฉ | Bright & Youthful Guide | 205 Hz | 1.00x | โถ๏ธ Preview |
shivani |
Shivani | ๐ฉ | Expressive Native Hindi | 215 Hz | 0.98x | โถ๏ธ Preview |
dev |
Dev | ๐จ | News Anchor & Tech Broadcast | 120 Hz | 1.00x | โถ๏ธ Preview |
sameer |
Sameer | ๐จ | Friendly Walkthrough Mentor | 125 Hz | 0.98x | โถ๏ธ Preview |
ravi |
Ravi | ๐จ | Deep Resonance Documentary Narrator | 112 Hz | 0.95x | โถ๏ธ Preview |
soham |
Soham | ๐จ | Soothing Classic Storyteller | 110 Hz | 0.92x | โถ๏ธ Preview |
๐ 3. Global & International Series
| Voice ID | Display Name | Gender | Persona Character | Pitch | Speed | Preview |
|---|---|---|---|---|---|---|
heart |
Heart | ๐ฉ | Global Flagship American English | 190 Hz | 1.00x | โถ๏ธ Preview |
bella |
Bella | ๐ฉ | Warm Engaging American Storyteller | 198 Hz | 1.00x | โถ๏ธ Preview |
sarah |
Sarah | ๐ฉ | Articulate Professional Presenter | 188 Hz | 1.00x | โถ๏ธ Preview |
adam |
Adam | ๐จ | Confident Commercial Voiceover | 118 Hz | 1.00x | โถ๏ธ Preview |
michael |
Michael | ๐จ | Natural Conversational Baritone | 114 Hz | 1.00x | โถ๏ธ Preview |
๐ Supported Languages & G2P Routing
Resona provides universal phonemization and neural voice rendering across 10 global and regional languages:
| Language | Code | G2P Engine | Default Flagship Voice |
|---|---|---|---|
| Hinglish | hinglish |
Resona Hybrid Transducer + IPA | anjura |
| Hindi | hi / hindi |
Devanagari eSpeak-NG IPA | meera |
| English (US) | en / en-us |
American English IPA | heart / arjun |
| English (UK) | en-gb |
British English IPA | heart |
| Spanish | es |
Spanish G2P IPA | heart |
| French | fr |
French G2P IPA | heart |
| Italian | it |
Italian G2P IPA | heart |
| Portuguese | pt |
Portuguese G2P IPA | heart |
| Japanese | ja |
Romaji / Kana IPA | heart |
| Mandarin Chinese | zh / cmn |
Pinyin IPA | heart |
โก Benchmarks & Performance
Measured on commodity consumer hardware (Intel Core i7-12700H @ CPU, single thread):
| Metric | Cloud APIs | Resona 82M | Advantage |
|---|---|---|---|
| Cold Start Latency | 800 - 1500ms | 180ms | ๐ข 4.4x Faster |
| Real-Time Factor (RTF on CPU) | Network Dependent | 0.18x | ๐ข 5.5x Realtime |
| Words Per Second (WPS) | ~15 words/s | 48 words/s | ๐ข 3.2x Faster |
| Air-Gapped / Zero Internet Required | โ No | โ Yes | ๐ข 100% Offline |
| Memory Footprint (RAM) | N/A | ~420 MB | ๐ข Ultra-Light |
Speed Rating
RTF 0.18x โโโโโโโโโโโโโโโโโโโโโโโโ 82% faster than real-time
๐ Repository Structure
Resona/
โโโ resona/ # 100% Self-Contained Python package
โ โโโ core/ # Neural vocoder, ResonaModel, ResonaPipeline, CustomSTFT
โ โโโ preprocessing/ # Normalizer, ResonaHinglishEngine, G2P phonemizer
โ โโโ voices/ # VoiceManager, registry.json, 21 style tensors & samples
โ โโโ vocab/ # Bundled lexicons, compound phrases & tech whitelist
โ โโโ cli.py # resona-tts CLI command
โโโ models/ # Checkpoints (resona-indic-v1.pth, resona-v1.pth, config.json)
โโโ vocab/ # Dictionaries: compound_phrases.json, hinglish_lexicon.json
โโโ examples/ # Working python examples (01_quickstart.py, 02_hinglish_tutorial.py)
โโโ tests/ # Automated test suite (model loading, voice registry, normalization)
โโโ docs/ # Detailed Voice Catalog & specifications (VOICES.md)
โโโ pyproject.toml # Build & packaging specifications (wheel + sdist)
โโโ checksums.sha256 # SHA-256 release integrity hashes
๐ License & Attribution
Resona is distributed under the Apache License 2.0. Upstream architectural foundations and legal notices are detailed in NOTICE.
Built with โค๏ธ for offline, privacy-first, and expressive speech synthesis.