Projects

Building Speech AI for Twi-Language Healthcare: The HARMLET Collaboration with Aarhus University

“Me ti yɛ me ya ɛna me temperature nso ɛyɛ high.”

Sentences like this — Twi carrying the grammar, English carrying the medical vocabulary — are how clinical conversations actually sound across much of Ghana. They are also what contemporary speech recognition handles worst. HARMLET, a research collaboration between Aarhus University and Releev AI, exists to change that.

The project

Supported by seed funding from iClimate, Aarhus University’s interdisciplinary climate research centre, HARMLET is developing AI-assisted telemedicine that lets patients in rural Ghanaian communities speak naturally, in their own language, to reach care that may otherwise be hours of travel away. The work runs through 2027 across three rural communities, beginning with Twi — spoken by more than eight million people — and extending to Ewe, Dagbani and Ga. Aarhus University leads the research programme; Releev AI contributes the machine learning and systems engineering.

The research gap

Preparatory work quantified three obstacles. First, model coverage: OpenAI’s Whisper, the most widely deployed speech recognition model, contains no Akan/Twi language token — transcription requests fail outright. Second, data asymmetry — and a domain gap beneath it. Twi is better served for synthesis than recognition: BibleTTS provides roughly 86 hours of studio-quality single-speaker audio, well suited to text-to-speech, where consistency is the goal. Recognition needs the opposite — many speakers, many accents, real acoustic conditions — and there the public record is thin. But the deeper gap is domain: what exists is scripture and general-purpose read speech. There is no clinical corpus for Twi, or for most low-resource languages: no drug names, no dosage instructions, no symptom descriptions, none of the structure of a real consultation. A model can be fluent in a language and still fail in a clinic. Building that corpus, with health workers and the communities the system serves, is the substance of HARMLET’s data work. Third, code-switching: mixing Twi and English within a sentence is the norm in clinical speech, not an edge case, and conventional word-error-rate metrics understate failures at precisely the switch points where drug names and dosages occur.

The approach

Rather than extending English-centric models, HARMLET builds on Meta’s open Massively Multilingual Speech (MMS) model, whose pretraining includes Twi. On this shared base, each language is served by a compact adapter (~60 MB) trained with parameter-efficient fine-tuning — so extending the platform to a new language requires a corpus and an adapter, not a new system.

The clinical corpus is co-created with the communities it serves: health workers develop the medical vocabulary and dialogue scenarios, contributors are compensated, and every recording is linked to a consent registry in which a withdrawal automatically removes that speaker’s data from all future training versions. Transcripts carry word-level language annotation, and evaluation includes a dedicated switch-point error metric alongside standard measures.

Environmental accountability

As an iClimate project, HARMLET treats its own footprint as a research question. Every training run logs its emissions, and the serving architecture is designed to meter carbon per request. The relevant comparison is the journey a remote consultation replaces: preliminary estimates place a voice consultation at a small fraction of a kilogram of CO₂-equivalent, against roughly 1.5–2.5 kg for a typical motorbike round trip to a clinic. Measured figures will be published as deployment matures.

Human oversight

HARMLET does not diagnose, and its output is never acted on autonomously. Transcripts support health workers within a staged deployment process — accuracy thresholds, clinician review, shadow operation, and rollback on regression — reflecting the reality that low-resource speech recognition will err, and that system design determines whether errors are caught by people or passed onward.

Current status

A first Twi text-to-speech model has been trained on BibleTTS and awaits evaluation with native speakers. Recognition training follows the clinical corpus now being collected. Integration of the full voice loop and field testing in the Living Lab communities follow from there.

HARMLET has the potential to transform health care delivery in rural and under-served communities and we are excited to be collaborating with Releev AI and leveraging their expertise in GenAI – Dugasseh Akowuge Frank (PhD), Aarhus University.

We will report results as they arrive, including the negative ones.

Releev AI builds AI-powered telemedicine and clinical decision support that extends hospital care into rural and peri-urban communities, with speech technology for under-resourced languages as its access layer.

HARMLET is a research collaboration with Aarhus University supported by iClimate seed funding.

Author

admin

Comment (1)

  1. One Sentence, Two Languages: Code-Switching in Clinical Speech AI – Releeve AI
    August 25, 2026

    […] Part 2 of our series on HARMLET, our healthcare speech-AI collaboration with Aarhus University. [Part 1] […]

Leave a comment

Your email address will not be published. Required fields are marked *