Distraction Free Reading

Automating Language, Losing Relationality: Feminist Reflections on African Language AI

A grandmother pauses before answering a child. The meaning of her words lies not only in the words themselves, but in tone, shared history, the relationship between speaker and listener, the setting, and even the silence that precedes the response. None of these become part of the training data for today’s large language models (LLMs). As Artificial Intelligence (AI) expands into African languages, discussions often celebrate new datasets, multilingual benchmarks, and translation systems. Yet something more fundamental is at stake. AI does not simply automate language. It transforms what counts as language into a textually legible dataset and abstracts language from the context of its production. A Large Language Model (LLM) is an AI system that has learned patterns from large amounts of words. This allows it to understand what people write or say and generate human-like responses, making it useful for tasks such as answering questions, drafting documents, translating languages, and assisting with problem-solving.

This blog entry emerges from an ongoing conversation between a South African computational linguist who is a first-language Sesotho speaker, born into a Setswana-speaking family, but whose competence in Setswana is that of a heritage speaker, and a linguistic anthropologist of African language activism. Our concerns intersect over the question of how communities of speakers orient to and are affected by digital language technologies. The rapid rise of neural-network-based artificial intelligence has placed African language development at the center of conversations about digital inclusion. Over the last several years, the continent has seen an expansion of initiatives, research networks, and community-driven projects—from Deep Learning Indaba and Masakhane to regional collectives such as GhanaNLP and EqualyzAI. These efforts, often supported by major tech companies and multi-stakeholder development schemes like the Lacuna Fund, are grounded in the promise that Large Language Models (LLMs) can open access to information for speakers of African languages that have historically been marginalized within global computational infrastructures.

LLM architectures now underlie translation tools, dialog systems, sentiment analysis pipelines, and even healthcare apps such as UlizaLlama, designed for Swahili-speaking women seeking medical information. A growing number of initiatives are developing AI-powered business and linguistic services for African languages, including Sesotho and isiZulu.

An image representing tones of Sesotho speech in waveform

Manual tonal annotation in Praat for Sesotho speech: The annotation was performed manually by an expert annotator, who assigned tonal labels based on a combination of auditory judgement and acoustic evidence. The waveform and spectrogram provide the speech signal, while the superimposed fundamental frequency (F0) contour (blue line) is used to verify tonal movement and alignment. Because tone in Sesotho is associated with the syllable rather than the individual segment, the annotation includes a dedicated syllable tier, with each syllable assigned a tonal category (e.g., High (H) or Low (L)). Additional tiers represent the phonetic segmentation and orthographic word boundaries, allowing tonal events to be interpreted within their phonological and lexical context. The highlighted region shows the currently selected syllable during the annotation process.

The overarching aim is to better serve communities whose languages have long been labeled “low-resource,” a category that obscures the complex political and ideological histories behind why these languages lack documented datasets in the first place. Indeed, African languages are far less represented in digital corpora than their European counterparts. This scarcity fuels well-known initiatives such as the African Dataset Challenge (Siminyu et al. 2020), but it also reinforces a familiar colonial narrative: Africa as a blank space on the map, awaiting “civilizing” by external actors. In treating datasets as cultural repositories, LLM development risks reproducing Enlightenment-era assumptions about language as a fixed, referential system (Schneider 2022; cf. Briggs and Baumann 2001) rather than a relational, embodied practice. The idea that language can be extracted, codified, and plotted on a linguistic map may suit machine learning pipelines, but it fails to acknowledge how meaning emerges through lived context, interpersonal relations, and social obligations.

This tension becomes clear when we consider languages shaped by deeply oral traditions. Many African languages encode meaning through tone, relational deixis, honorifics, and context-dependent narrative forms. In Sesotho or isiZulu, for example, one cannot separate what is said from who says it, in what setting, and under what moral or ritual conditions (Fandrych 2012). Yet LLMs flatten language into statistical patterns (McCoy et al. 2019), stripping away metadata such as tone, prosody, and speaker roles. As Bender and Koller (2020) argue, these systems become “syntactic engines” with limited access to communicative intent.

Tone marking illustrates the problem starkly sincetone is vital for distinguishing lexical and grammatical meaning in most Bantu languages. However, in an earlier era of standardization,(post)-colonial linguists (Errington 2001) developed orthographic systems that did not mark tone in writing. As a result, LLMs trained primarily on written texts, often drawn from religious or encyclopedic sources, fail to capture tonal contrasts in structures such as narrative mood, cautionary registers, or speculative storytelling, forms of oral expression such as honorifics in Sesotho or Zulu (Nurse 2008; Irvine 2022).

The remarkable success of large language models is built on a simple premise: linguistic competence can be approximated by learning statistical regularities from massive collections of text. This logic has made standardization indispensable, enabling the aggregation of large, consistent corpora for model training. Yet standardization is not a neutral process. While it makes language computationally tractable, it also privileges those forms of language that are written, codified, and widely available, often excluding the rich oral, contextual, and performative dimensions through which meaning is created in many African languages. This is not merely a technical omission. It represents a cultural erasure that privileges the written over the oral, the standardized over the vernacular, and the universal over the situated (Bird 2020).

Such losses also affect relational personhood – the idea, central to many African philosophies, that individuals are constituted through their involvement in social and communicative networks. Language in this worldview is not a neutral tool but an enactment of linguistic ideologies for maintaining obligations, histories, and moral hierarchies. When AI systems fail to capture these relational forms, they risk misrepresenting not just the language, but the social worlds it sustains, the way a grandmother forges a kinship with her grandchild, for example.

This perspective aligns closely with the work of African technologists who highlight how multilingualism, linguistic complexity, and tone challenge English-oriented NLP assumptions (Nekoto et al. 2020). In datasets like SAfriSenti, for instance, more than a third of annotated tweets involve code-switching and reveal the fluid, improvisational nature of everyday communication (Mabokela and Schlippe, 2022). Practices such as code-switching and multimodal expression point toward more expansive understandings of personhood and communicative agency. A feminist approach therefore urges AI practitioners to move beyond text-only corpora and embrace multimodal data: tonal speech, gesture, storytelling, music, and culturally specific metadata. Emerging projects, such as the Sotho-Tswana Multimodal Music Information Retrieval dataset, show the promise of integrating audio, visual, and textual modalities. Yet even these efforts will be challenged to embrace local epistemologies, allowing native speakers to annotate not only linguistic categories but cultural meanings.

Ultimately, a more relational, multimodal, and community-centered AI would recognize that incompleteness (Nyamnjoh 2017)—rather than striving for totalizing linguistic capture—can be an ethical stance. Rather than importing “universal” ethics frameworks—fairness, explainability, trust—from Euro-American computer science, African communities should be able to define ethical commitments grounded in lived language use. Feminist and Post-colonial approaches to Science and Technology (Harding 2009) are complicit with the move to decenter such frameworks, urging us to value relationality over representation (Metz 2012), and communicative practice over abstract code If AI is to engage meaningfully with African languages, it must learn not just from language data, but with the social worlds that sustain it.

The future of African language AI is therefore not simply a question of supporting more languages or building larger models. It is a question of deciding what kind of language—and ultimately what kind of relationships—intelligent systems could possibly sustain. If AI learns only co-occurring words while forgetting the relations that brought those words together in the first place, then digital inclusion will remain partial and elusive African language technologies will succeed not when they merely speak more languages, but when they remain accountable to the social worlds that give those languages meaning.

This post is part of a series on an African Studies approach to techno-ethics and feminist AI, read the introduction to this series here.


This post was curated by Contributing Editor Volney Friedrich and reviewed by  Tayeba Batool

References

Aiseng, Kealeboga. 2026. “Decolonizing Algorithms: AI-Driven Media and the Reclamation of African Linguistic Identities.” Emerging Media 4 (2): 376–94. https://doi.org/10.1177/27523543261447913.

Bauman, Richard, and Charles L. Briggs, eds. 2003. Voices of Modernity: Language Ideologies and the Politics of Inequality. Studies in the Social and Cultural Foundations of Language 21. Cambridge University Press.

Bender, Emily M., and Alexander Koller. 2020. “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data.” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–98. https://doi.org/10.18653/v1/2020.acl-main.463.

Bird, S. (2020). Decolonising Speech and Language Technology. Proceedings of the 28th International Conference on Computational Linguistics, 3504–3519.

Errington, Joseph. 2001. “Colonial Linguistics.” Annual Review of Anthropology 30: 19–39.

Fandrych, I., 2012. Between tradition and the requirements of modern life: Hlonipha in southern Bantu societies, with special reference to Lesotho. Journal of Language and Culture, 3(4), pp.67-73. https://doi.org//10.5897/JLC11.056 

Harding, Sandra. 2009. “Postcolonial and Feminist Philosophies of Science and Technology: Convergences and Dissonances.” Postcolonial Studies 12 (4): 401–21. https://doi.org/10.1080/13688790903350658.

Irvine, Judith T. 2022. “Ideologies of Honorific Language.” Pragmatics. Quarterly Publication of the International Pragmatics Association (IPrA), July 6, 251–62. https://doi.org/10.1075/prag.2.3.02irv.

McCoy, R.T., Pavlick, E. and Linzen, T., 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. arXiv preprint arXiv:1902.01007. https://doi.org/10.48550/arXiv.1902.01007

Metz, T. 2012. “An African Theory of Moral Status: A Relational Alternative to Individualism and Holism.” Ethical Theory and Moral Practice 15 (3): 387–402. https://doi.org/10.1007/s10677-011-9302-y.

Nekoto, Wilhelmina, Vukosi Marivate, Tshinondiwa Matsila, et al. 2020. “Participatory Research for Low-Resourced Machine Translation: A Case Study in African Languages.” In Findings of the Association for Computational Linguistics: EMNLP 2020, edited by Trevor Cohn, Yulan He, and Yang Liu. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.195.

Nurse, Derek. 2008. Tense and Aspect in Bantu. OUP Oxford.

Nyamnjoh, F. 2017. “Incompleteness: Frontier Africa and the Currency of Conviviality.” Journal of Asian and African Studies 52 (3): 253-270. https://doi.org/10.1177/0021909615580867

Schneider, Britta. 2022. “Multilingualism and AI: The Regimentation of Language in the Age of Digital Capitalism.” Signs and Society 10 (3): 362–87. https://doi.org/10.1086/721757.

Siminyu, Kathleen, Sackey Freshia, Jade Abbott, and Vukosi Marivate. 2020. “AI4D — African Language Dataset Challenge.” arXiv:2007.11865. Preprint, arXiv, July 23. http://arxiv.org/abs/2007.11865.Ronny Mabokela and Tim Schlippe. 2022. A Sentiment Corpus for South African Under-Resourced Languages in a Multilingual Context. In Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages, pages 70–77, Marseille, France. European Language Resources Association.

Leave a Reply

Your email address will not be published. Required fields are marked *