Large language models are trained on text that reflects a particular cultural gaze. When those models are then used to represent Indigenous communities — without consultation, without consent, without the underlying ontology shifting — the result is exoticization, misappropriation, and erasure dressed up in fluent prose.
My forthcoming chapter, AI Techniques for Indigenous Cultural Expression, argues that the fix is not to add Indigenous content to existing systems. It is to design systems that operate according to Indigenous ways of knowing in the first place.
The chapter offers five practical methods:
- Culturally informed prompting. Establish cultural context explicitly. Incorporate Indigenous terms and concepts. Test prompts with community members, refine, document.
- Ethical data collection. Apply the CARE principles — collective benefit, authority to control, responsibility, ethics — to every dataset that touches Indigenous knowledge.
- Retrieval-Augmented Generation for contextual sensitivity. Ground generated outputs in ethically curated sources the community continues to own. Keep provenance legible.
- Fine-tuning for cultural integrity. Treat human variation as signal, not noise.
- Coalition-based coordination using multi-agent systems. Replace single-model authority with coordination protocols that let multiple knowledge frameworks coexist.
These methods rest on four ontological orientations — pluralism, groundedness, liveliness, enactment — and are implemented through active inference. The theoretical work is laid out in the chapter.
One example, already public: Te Hiku Media’s te reo Māori speech recognition project. Starting in 2016, Te Hiku recorded native speakers, amassed 310 hours of community-owned speech-text pairs through their Whare Kōrero platform, and trained models that reached 92% accuracy on te reo and 82% on bilingual transcription. Contributors retained ownership throughout. The dataset is the model’s substrate, and the dataset stayed with the community.
What runs underneath all five methods is the same commitment: Indigenous communities own, control, and access the data that derives from them. Models trained on that data, and the protocols that govern it, are extensions of community sovereignty rather than mechanisms that route around it.
These methods aren’t theoretical for us. The voice-data pipeline on the Now page is one place they meet practice.
This essay is a pointer. The chapter is forthcoming in Envisioning Indigenous Methods in Digital Media and Ecologies, edited by Ashley Cordes.