Documentation
ልሳን (LISAN) — Offline-First Voice AI Languae translator for Ethiopian Languages.
Real-time, AI powered, on-device speech-to-speech translation for Ethiopian languages working almost on all budget phones.
PROJECT NAME
ልሳን (LISAN) — Offline-First Voice AI for Ethiopian Languages
INSPIRATION
Ethiopia is home to over 120 million people speaking more than 80 distinct languages. While this linguistic diversity is a profound cultural heritage, it creates acute communication barriers in everyday life:
• Healthcare and Emergency Access: In regional referral hospitals and rural clinics, doctors and nurses often do not speak the regional mother tongue of their patients (for example, Amharic-speaking medical personnel treating Afaan Oromo or Somali speakers). In emergency triage situations, miscommunication carries life-altering consequences.
• Rural Commerce and Trade: In open-air agricultural markets, grain and coffee traders travel between regional states where different languages intersect daily.
• Civic Access and Justice: Navigating administrative offices, transport hubs, and public services without language parity compromises human dignity and equal opportunity.
Every major translation app on the market today (Google Translate, cloud translation APIs) assumes a continuous, high-speed internet connection. In rural Ethiopia and across the Horn of Africa, frequent network blackouts, remote cellular dead zones, and steep mobile data costs render cloud tools useless when they are needed most.
Furthermore, Ethiopian languages have long suffered from "low-resource" status by global tech companies—plagued by broken tokenization, awkward literal syntax, and lack of speech support.
We asked: What if two people speaking completely different mother tongues could hold a single phone between them, press a single tactile button, and converse naturally in real-time — with 100% privacy, sub-6-second speed, and absolutely ZERO internet connection?
That vision became LISAN (ልሳን) — derived from the ancient Ge'ez root for "tongue", "language", and "human speech".
WHAT IT DOES
LISAN is a standalone, offline-capable mobile voice communicator. Two individuals who speak different languages hold one phone and communicate through it seamlessly.
• 100% Offline and On-Device: All speech recognition, neural machine translation, and speech synthesis execute locally on the smartphone's processor. Zero cloud packets, zero data costs, zero latency spikes, and complete data privacy.
• Natural Conversational Flow: Speaker A speaks in their native tongue; LISAN transcribes, translates, and vocalizes the sentence in Speaker B's language in seconds. Speaker B replies with a single tap, reversing the flow.
• Supported Languages:
-
Amharic (Ethiopic Ge'ez script)
-
Afaan Oromoo (Latin Qubee script)
-
Tigrinya (Ethiopic Ge'ez script)
-
Somali (Latin script)
-
English (for tourists, aid workers, and global visitors)
• Tactile Industrial Design: Built around a clean, tactile design language inspired by Dieter Rams and Teenage Engineering. High-contrast typography, warm sand and obsidian palettes, and responsive acoustic ripple rings indicate speech status clearly without requiring users to read dense text.
HOW WE BUILT IT: GOOGLE CLOUD TRAINING TO EDGE DEPLOYMENT
LISAN is engineered on a hybrid AI architecture: cloud-scale fine-tuning on GOOGLE COLAB followed by model compression and INT8 quantization for deployment on Flutter Mobile (Android and iOS).
1. Cloud Fine-Tuning with GOOGLE COLAB
• Compute Infrastructure
• Dataset: Trained across 64,623 training pairs and 3,401 validation pairs sourced from the open Ethiopian language corpus dataset.et (managed by Snapwre).
• QLoRA Fine-Tuning: Used PyTorch 2.x, Hugging Face PEFT, and BitsAndBytes 4-bit QLoRA to fine-tune Meta NLLB-200 Distilled 600M. This reduced training VRAM from 18GB down to under 8GB, enabling rapid hyperparameter sweeps on cost-effective AWS instances in approximately 2 hours per epoch.
2. Edge Compression and ONNX Quantization
• Merged the LoRA adapters with the base transformer and exported the sequence-to-sequence graph to ONNX.
• Applied symmetric INT8 dynamic quantization, compressing the model footprint to approximately 609 MB so it fits comfortably within the RAM limits of budget smartphones.
3. Sub-Second Speech-to-Text (STT)
• Deployed OpenAI Whisper Tiny locally on-device.
• Traditional Whisper implementations invoke external FFmpeg subprocesses and disk writes that add 1.5 to 2.0 seconds of latency. We re-engineered the audio capture pipeline to stream raw 16kHz audio directly into an in-memory NumPy buffer via soundfile, slashing transcription latency to 0.41 seconds.
4. Cross-Platform Mobile Client (Flutter)
• Built a native Flutter application with a custom C-interop bridge to the ONNX Runtime Mobile C-API.
• Implemented the custom LisanTheme design system with calibrated Ethiopic fidel typography and native offline speech synthesis.
CHALLENGES WE RAN INTO
1. Sub-6s Edge Latency on Modest Hardware: The translation model initially exhibited a 13-second "cold-start" delay on the first query, and beam search with num_beams=4 was sluggish on mobile CPUs. We engineered a background engine warm-up during app boot and tuned beam search to num_beams=2, cutting latency by 65% with zero perceptible quality loss (achieving a ~5.9s end-to-end roundtrip).
2. Tokenizer Code Clashes: Meta's NLLB-200 maps Afaan Oromoo to gaz_Latn (West Central Oromo) rather than the standard orm_Latn. Querying orm_Latn produced silent output failures. We engineered a strict normalization registry to correctly route regional language codes.
3. Ge'ez Script Encoding Hazards: Non-Latin Ethiopic characters require careful stream handling. We enforced UTF-8 byte stream validation across all audio, tokenizer, and UI buffers to eliminate glyph corruption.
4. Mobile Memory Budgets: Low-end smartphones kill apps exceeding 1GB RAM. Through INT8 quantization and memory-mapped model weights, LISAN maintains an active working set under 800MB RAM.
ACCOMPLISHMENTS THAT WE'RE PROUD OF
• Zero Cloud Latency, 100% Autonomy: Real-time neural voice translation running entirely on a mobile phone without internet connectivity.
• 0.41s Speech Recognition: Ultra-fast local speech transcription on edge hardware.
• Bridging Low-Resource Languages: Accurate bidirectional speech translation between Amharic, Afaan Oromo, Tigrinya, Somali, and English.
• Product-Grade Visual Identity: A warm, tactile industrial design system that avoids generic AI clichés and respects Ethiopian cultural heritage.
WHAT WE LEARNED
• GOOGLE COLAB Democratizes Low-Resource NLP: Foundation models can be adapted to underserved languages with accessible GPU resources in hours rather than months.
• Memory I/O is Key to Edge Speed: In on-device AI pipelines, audio format conversion and disk I/O often cause more lag than model inference itself. Moving to in-memory streaming was our single biggest latency win.
• Linguistic Architecture: Translating between Semitic languages (Amharic, Tigrinya) and Cushitic languages (Afaan Oromo, Somali) requires specialized tokenization due to fundamentally different morphological structures.
WHAT'S NEXT FOR LISAN
• Automated Pipelines: Establishing an automated community-in-the-loop pipeline to ingest and fine-tune regional dialects on AWS.
• Expanding Regional Languages: Adding support for Sidama, Wolaytta, Gurage, and Afar.
• Full-Duplex Simultaneous Translation: Transitioning from push-to-talk to continuous ambient translation using streaming CTC decoders.
