๐ŸŒŸ Why Resona?

Traditional TTS engines either require expensive cloud APIs with latency spikes or generate robotic, metallic voices when dealing with Indian languages and Hinglish code-switching. Resona solves this with an ultra-compact 82M parameter neural architecture running 100% offline on standard consumer CPUs.


๐Ÿ‡ฎ๐Ÿ‡ณ Native Bilingual Code-Switching

Seamlessly transitions between English UI / technical terminology and native Hindi diction without unnatural accent distortion or robotic stutter.

โšก Real-Time Edge Latency

Generates speech in sub-second time on standard commodity CPUs (RTF < 0.20x). Zero GPU or specialized hardware required.

๐Ÿซ Natural Breathing Cadence

Built-in biological rhythm injector that inserts subtle 180ms clause pauses and 350ms sentence boundaries for conversational realism.

๐Ÿ”’ 100% Private & Air-Gapped

Zero external API calls, zero telemetry, and zero data leakage. All weights, style vectors, and lexicons run locally on your device.


๐Ÿ“Š Pipeline Architecture

flowchart LR
    A[Raw Input Text] --> B[Text Normalizer]
    B --> C{Language Router}
    C -- Hinglish --> D[Hinglish Transducer]
    D --> E[Protected Tech Whitelist Guard]
    E --> F[Devanagari / IPA Phonemizer]
    C -- Multilingual / English --> F
    F --> G[Token & Rhythm Alignment]
    H[Voice Style Tensor 256-D] --> I[AdaIN Prosody Predictor]
    G --> I
    I --> J[82M Resona Neural Vocoder]
    J --> K[Natural Breathing Cadence]
    K --> L[Studio WAV Audio 24kHz]

    style A fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style D fill:#313244,stroke:#f38ba8,stroke-width:2px,color:#cdd6f4
    style E fill:#313244,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4
    style J fill:#45475a,stroke:#cba6f7,stroke-width:3px,color:#f5e0dc
    style L fill:#1e1e2e,stroke:#a6e3a1,stroke-width:2px,color:#a6e3a1

๐Ÿš€ Quick Start

1. Installation

# Clone the repository
git clone https://github.com/MilannSharma/Resona.git
cd Resona

# Install package dependencies
pip install -e .

Note: Requires espeak-ng installed on your host system (standard package on Linux/macOS, or via installer on Windows). Resona includes automatic multi-path discovery and CLI fallback.


2. Python API

from resona import ResonaPipeline

# 1. Initialize pipeline with flagship Hinglish voice (Anjura)
pipeline = ResonaPipeline(voice="anjura", language="hinglish")

# 2. Synthesize with natural code-switching & breathing cadence
text = "Resona audio engine se speech generate karna super fast aur natural hai. File menu par click kijiye."
result = pipeline.synthesize(text, output_path="output.wav", speed=0.95)

print(f"Generated {result.duration_seconds:.2f}s studio audio at {result.sample_rate}Hz!")

3. Terminal CLI (resona-tts)

# List all 21 available Grade A personas
resona-tts --list-voices

# Synthesize speech directly from command line
resona-tts --text "Welcome to the Resona Neural Audio Engine." \
           --voice arjun \
           --output welcome.wav

# Synthesize Hinglish tutorial with custom speed multiplier
resona-tts --text "Settings menu par click kijiye aur resolution 1080p select kijiye." \
           --voice anjura \
           --lang hinglish \
           --speed 0.95 \
           --output settings.wav

๐ŸŽ™๏ธ Studio Voice Roster (21 Personas)

All 21 voices include pre-rendered, high-fidelity reference .wav previews in resona/voices/samples/:

๐Ÿ‡ฎ๐Ÿ‡ณ 1. Hinglish & Indic Flagships (Studio Curated)

Voice ID Display Name Gender Persona Character Pitch Speed Preview
anjura Anjura ๐Ÿ‘ฉ Flagship Corporate Explainer 184 Hz 0.95x โ–ถ๏ธ Preview
divya Divya ๐Ÿ‘ฉ Deep, Articulate Explainer 184.6 Hz 0.95x โ–ถ๏ธ Preview
meera Meera ๐Ÿ‘ฉ Warm Conversational Native Hindi 210 Hz 1.00x โ–ถ๏ธ Preview
priya Priya ๐Ÿ‘ฉ Dynamic Tech Educator 195 Hz 1.00x โ–ถ๏ธ Preview
arjun Arjun ๐Ÿ‘จ Enterprise Baritone Narrator 115 Hz 0.95x โ–ถ๏ธ Preview
kabir Kabir ๐Ÿ‘จ Executive Conversational Podcast 128 Hz 0.95x โ–ถ๏ธ Preview
aman Aman ๐Ÿ‘จ Agile Tech Walkthrough (Denoised) 132 Hz 1.00x โ–ถ๏ธ Preview
atul Atul ๐Ÿ‘จ Classic Indic Corporate Voice 122 Hz 0.98x โ–ถ๏ธ Preview

๐ŸŽ™๏ธ 2. Indic English & Native Hindi Series

Voice ID Display Name Gender Persona Character Pitch Speed Preview
ananya Ananya ๐Ÿ‘ฉ Friendly Academic Explainer 190 Hz 1.00x โ–ถ๏ธ Preview
nisha Nisha ๐Ÿ‘ฉ Calm, Gentle Corporate Presenter 182 Hz 0.98x โ–ถ๏ธ Preview
tara Tara ๐Ÿ‘ฉ Bright & Youthful Guide 205 Hz 1.00x โ–ถ๏ธ Preview
shivani Shivani ๐Ÿ‘ฉ Expressive Native Hindi 215 Hz 0.98x โ–ถ๏ธ Preview
dev Dev ๐Ÿ‘จ News Anchor & Tech Broadcast 120 Hz 1.00x โ–ถ๏ธ Preview
sameer Sameer ๐Ÿ‘จ Friendly Walkthrough Mentor 125 Hz 0.98x โ–ถ๏ธ Preview
ravi Ravi ๐Ÿ‘จ Deep Resonance Documentary Narrator 112 Hz 0.95x โ–ถ๏ธ Preview
soham Soham ๐Ÿ‘จ Soothing Classic Storyteller 110 Hz 0.92x โ–ถ๏ธ Preview

๐ŸŒ 3. Global & International Series

Voice ID Display Name Gender Persona Character Pitch Speed Preview
heart Heart ๐Ÿ‘ฉ Global Flagship American English 190 Hz 1.00x โ–ถ๏ธ Preview
bella Bella ๐Ÿ‘ฉ Warm Engaging American Storyteller 198 Hz 1.00x โ–ถ๏ธ Preview
sarah Sarah ๐Ÿ‘ฉ Articulate Professional Presenter 188 Hz 1.00x โ–ถ๏ธ Preview
adam Adam ๐Ÿ‘จ Confident Commercial Voiceover 118 Hz 1.00x โ–ถ๏ธ Preview
michael Michael ๐Ÿ‘จ Natural Conversational Baritone 114 Hz 1.00x โ–ถ๏ธ Preview

๐ŸŒ Supported Languages & G2P Routing

Resona provides universal phonemization and neural voice rendering across 10 global and regional languages:

Language Code G2P Engine Default Flagship Voice
Hinglish hinglish Resona Hybrid Transducer + IPA anjura
Hindi hi / hindi Devanagari eSpeak-NG IPA meera
English (US) en / en-us American English IPA heart / arjun
English (UK) en-gb British English IPA heart
Spanish es Spanish G2P IPA heart
French fr French G2P IPA heart
Italian it Italian G2P IPA heart
Portuguese pt Portuguese G2P IPA heart
Japanese ja Romaji / Kana IPA heart
Mandarin Chinese zh / cmn Pinyin IPA heart

โšก Benchmarks & Performance

Measured on commodity consumer hardware (Intel Core i7-12700H @ CPU, single thread):

Metric Cloud APIs Resona 82M Advantage
Cold Start Latency 800 - 1500ms 180ms ๐ŸŸข 4.4x Faster
Real-Time Factor (RTF on CPU) Network Dependent 0.18x ๐ŸŸข 5.5x Realtime
Words Per Second (WPS) ~15 words/s 48 words/s ๐ŸŸข 3.2x Faster
Air-Gapped / Zero Internet Required โŒ No โœ… Yes ๐ŸŸข 100% Offline
Memory Footprint (RAM) N/A ~420 MB ๐ŸŸข Ultra-Light

Speed Rating

RTF 0.18x โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘ 82% faster than real-time


๐Ÿ“ Repository Structure

Resona/
โ”œโ”€โ”€ resona/                     # 100% Self-Contained Python package
โ”‚   โ”œโ”€โ”€ core/                   # Neural vocoder, ResonaModel, ResonaPipeline, CustomSTFT
โ”‚   โ”œโ”€โ”€ preprocessing/          # Normalizer, ResonaHinglishEngine, G2P phonemizer
โ”‚   โ”œโ”€โ”€ voices/                 # VoiceManager, registry.json, 21 style tensors & samples
โ”‚   โ”œโ”€โ”€ vocab/                  # Bundled lexicons, compound phrases & tech whitelist
โ”‚   โ””โ”€โ”€ cli.py                  # resona-tts CLI command
โ”œโ”€โ”€ models/                     # Checkpoints (resona-indic-v1.pth, resona-v1.pth, config.json)
โ”œโ”€โ”€ vocab/                      # Dictionaries: compound_phrases.json, hinglish_lexicon.json
โ”œโ”€โ”€ examples/                   # Working python examples (01_quickstart.py, 02_hinglish_tutorial.py)
โ”œโ”€โ”€ tests/                      # Automated test suite (model loading, voice registry, normalization)
โ”œโ”€โ”€ docs/                       # Detailed Voice Catalog & specifications (VOICES.md)
โ”œโ”€โ”€ pyproject.toml              # Build & packaging specifications (wheel + sdist)
โ””โ”€โ”€ checksums.sha256            # SHA-256 release integrity hashes

๐Ÿ“„ License & Attribution

Resona is distributed under the Apache License 2.0. Upstream architectural foundations and legal notices are detailed in NOTICE.

Built with โค๏ธ for offline, privacy-first, and expressive speech synthesis.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support