Speaker 1: Hey Mason! July ninth, twenty twenty-six. We have a packed show today, starting with a quick radar briefing on the newest research we have been keeping an eye on, followed by our weekly Native language technology roundup. Speaker 2: I am looking forward to that roundup, especially with all the recent movement in community-governed data models. Speaker 1: It is moving fast, but let us start with the research first. Speaker 2: Right, what is hitting the radar this week? Speaker 1: Mason, the research filter has surfaced a few developments that feel particularly relevant to our work on language sovereignty and reliable pipelines. Speaker 2: I am ready, let's see if there is anything substantial. Speaker 1: There is a strong theme emerging around how we bridge the gap for low-resource languages through transfer learning and cross-lingual reasoning. Speaker 2: Are we talking about just more data, or actual architectural shifts? Speaker 1: It is more about the mechanics of transfer. For instance, Lukmal Ilyas and Nevidu Jayatilleke have been looking at cross-lingual transfer learning for speech recognition, moving from Sinhala to Dhivehi. Speaker 2: That sounds useful for us, especially if we can apply those same transfer principles to Lakota speech models without needing massive new datasets. Speaker 1: Exactly, and it ties into a larger study by Andrea Alfarano and several others, which found that improving reasoning capabilities in English actually boosts performance in low-resource languages. Speaker 2: That is a vital distinction. It suggests that if we strengthen the underlying logic of the model, the linguistic application follows more naturally. Speaker 1: It also highlights the need for better evaluation. Daryna Dementieva and her team are working on PluraMath, which tries to extend mathematical reasoning evaluations beyond just high-resource languages. Speaker 2: We need that. If we are building tools for the community, we cannot just assume the model is reasoning correctly in Lakota just because it works in English. Speaker 1: Right, and that brings us to the problem of complex scripts. Burte Bayarsaikhan and colleagues are researching cognitive pivot translation for Mongolian, specifically dealing with the challenges of traditional scripts and digraphic systems. Speaker 2: That is a huge hurdle for us with legacy texts. Dealing with script-induced ambiguity is a nightmare for any optical character recognition system. Speaker 1: It really is. And speaking of the intersection of culture and technology, Aparna Madva and her co-authors have a paper on rethinking artificial intelligence through the lens of cultural heritage preservation. Speaker 2: That sounds less like a technical manual and more like a framework for how we should be approaching this entire project. Speaker 1: It is. It addresses how linguistic diversity and cultural preservation must drive the development of artificial intelligence for underrepresented groups. Speaker 2: It is a good reminder that our technical choices are actually political ones. Speaker 1: They are. Finally, on the technical side of our retrieval-augmented generation work, Felix Feldman and his team published a study on evaluating retrieval quality and faithfulness in public health question answering. Speaker 2: That is the core of what we are doing with the community-approved books. If the retrieval is poor or the model hallucinates, the whole system loses trust. Speaker 1: Precisely. Their systematic approach to evaluating how faithful a model is to its retrieved sources is something we should definitely look at for our own evaluation stages. Speaker 1: We are looking at a week where the tension between massive corporate infrastructure and local community control is becoming impossible to ignore, Mason. Speaker 2: It feels like we are seeing two different worlds colliding, where one side is trying to fix mistakes and the other is trying to build entirely new foundations. Speaker 1: Exactly, and that starts with a massive win for pronunciation in New Zealand, where Google Maps has finally rolled out a new voice to correctly pronounce Te Reo Maori place names. Speaker 2: Wait, so Google actually listened to the complaints about them butchering the names of towns and cities? Speaker 1: They did, but they didn't do it alone. They partnered with Te Taura Whiri i te Reo Maori, which is the Maori Language Commission, to train the voice on a local speaker over about six months. Speaker 2: That sounds better than just scraping random audio from the internet, but is it really sovereignty if Google still owns the platform? Speaker 1: That is the million dollar question. The chief executive of the commission, Ngahiwi Apanui-Barr, called it a step toward securing the language in the digital age, but it is still a tool within a corporate ecosystem. Speaker 2: It is a functional improvement, certainly, but it feels like a patch on a system that was originally built without them in mind. Speaker 1: It is a patch, but it sets a precedent for how these giants might have to work with language commissions moving forward. It is a far cry from the situation in South Asia, which we saw discussed at a United Nations side event this week. Speaker 2: I saw that. They were talking about the difference between participation and actual power, right? Speaker 1: Precisely. The discussion focused on whether nations in South Asia are building real capacity or just a more sophisticated kind of dependency on foreign tech. Speaker 2: And the concern there is that Indigenous ecological knowledge and oral traditions are being sucked into these systems without any clear consent or accountability. Speaker 1: Right, even when there are international obligations like the United Nations Declaration on the Rights of Indigenous Peoples, there is no clear actor to hold responsible when tribal data is harvested. Speaker 2: It sounds like a digital version of the old extractive models, where the resources are taken, processed elsewhere, and the original owners see none of the profit or control. Speaker 1: It really highlights why we cannot just talk about ethics as an add on. There was a piece published during the National Aboriginal and Torres Strait Islander Observance Week in Australia arguing that Artificial Intelligence must be built with Indigenous Knowledges, not against them. Speaker 2: They used a medical example, didn't they? Something about diabetic retinopathy screening? Speaker 1: Yes, they pointed to Aboriginal medical clinics in Western Australia that are trialing Artificial Intelligence to help screen for eye disease. That is being cited as a model of how it should work because it is grounded in the community. Speaker 2: So, instead of the technology being dropped into a community from the top down, the community is part of the design and the implementation. Speaker 1: Exactly. The argument is that we need to move from ethical guidelines to mandatory foundations like Indigenous Data Sovereignty, using frameworks like the CARE principles or the OCAP principles. Speaker 2: That brings us to the bigger structural fight. I was reading an analysis about the three different types of sovereignty at play in the Artificial Intelligence era: state, corporate, and Indigenous techno-sovereignty. Speaker 1: That is a heavy concept, but it is vital for our work with Lakota. The traditional idea of state sovereignty just does not work when the actual power flows through corporate-controlled servers and infrastructure. Speaker 2: So, even if a nation-state claims to protect its people, if the data is sitting in a private data center halfway across the world, who is actually in charge? Speaker 1: That is why the analysis points to groups like Te Hiku Media in New Zealand. They built a ninety-two percent accurate speech recognition model for Te Reo Maori themselves. They are not waiting for Google to fix their pronunciation; they are building the engine. Speaker 2: That is the difference between being a user and being a builder. Speaker 1: It is the difference between being a subject and being a sovereign. And we see a more personal, grassroots version of this in the United States, with a Navajo man named George R. Joe. Speaker 2: He built that mobile app, right? Tribal Trailz? Speaker 1: Yes, he is a former educator who spent three years developing a GPS-activated audio tour app. It covers routes like Gallup to Flagstaff and Flagstaff to Phoenix. Speaker 2: So, as you drive, it narrates the history of the land through a Navajo, Zuni, and Acoma lens? Speaker 1: Exactly. It is designed to correct the misconceptions that tourists often have about Native peoples and the lands they are driving through. Speaker 2: I love that. It is using the very technology that often erases culture to actually re-insert it into the landscape. Speaker 1: It is a way of reclaiming the narrative of the road. He is already moving into phase two, adding segments for Albuquerque and Santa Fe. Speaker 2: It is a small-scale version of what we are trying to do, but it is incredibly powerful because it is controlled entirely by the community. Speaker 1: It connects everything we have been talking about. Whether it is Google fixing a voice in New Zealand, or a Navajo educator building a tour app, or Australian clinics using medical AI, the core issue is the same. Speaker 2: It is about who holds the keys to the data and who gets to decide how the machine interprets the world. Speaker 1: And for us, as we look at Lakota, it means we cannot just wait for the big players to decide how our language sounds or how our history is told. Speaker 2: We have to be the ones building the models, the ones governing the data, and the ones ensuring that the technology serves the sovereignty of the people, not just the interests of the corporations. Speaker 1: That covers our look at everything from the new Māori voice on Google Maps to the critical fight for data sovereignty in South Asia. Speaker 2: It is a lot to process, but it really highlights how much is at stake when technology meets traditional knowledge. Speaker 1: Exactly, and it is a conversation that is only just beginning. That is all for our weekly roundup of Native research and language technology, Mason. We will see you next time.