Hey Mason! It's July ninth, twenty twenty-six, and there's a good one on deck today. Two stops ahead: first, a radar briefing on the newest research we've kept off the pile — the papers and preprints that actually earned a place in the pipeline this week, not just the ones that showed up. Then we swing into the weekly roundup on Native language tech, where the sovereignty questions keep getting sharper and the tooling keeps getting more interesting. Community-approved data, per-stage evaluation, the whole slow, careful build — there's real movement to talk about. So let's get into it. Coffee's hot, road's long, and the radar's picked up a few signals worth slowing down for. On the RAG side, there's a new systematic evaluation of retrieval-augmented generation for public health question answering, from Felix Feldman and colleagues. It walks through different retrieval configurations and measures faithfulness, basically how well the generated answer actually reflects what was retrieved rather than drifting off into invented content. That's squarely useful groundwork for anyone building an evaluation harness for a RAG pipeline, since public health answering shares the same core problem as grounding answers in community-approved Lakota texts: you need the system citing real source material, not hallucinating around it. A second cluster this week is about what it means to build AI for languages and cultures that mainstream NLP has ignored. One paper reframes Indic language AI through a lens of cultural heritage preservation rather than pure benchmark performance, arguing that linguistic diversity and cultural context need to shape how these systems get built, not just what they're scored on. That framing maps closely onto language sovereignty work: the goal isn't a benchmark number, it's a system a community actually trusts and controls. There's also a cross-lingual transfer cluster worth flagging together. One paper takes speech recognition from Sinhala to Dhivehi, showing how a model trained on a better-resourced related language can be adapted efficiently to a much lower-resource one. That's directly relevant to data-efficient adaptation strategies, the same kind of thinking that would apply to bootstrapping speech or text models for Lakota from related resources rather than starting from nothing. A related paper looks at multilingual multiple-choice reasoning and finds that having a model reason in English first, then answer in a low-resource language, improves accuracy and gives a better handle on the model's uncertainty. That's a concrete, testable idea for a per-stage evaluation pipeline: measure whether reasoning-then-translating actually helps reliability for underserved languages, rather than assuming an end-to-end approach in the low-resource language is automatically better. Then two more specific but still relevant papers. One extends math reasoning evaluation beyond high-resource languages, essentially pointing out that most reasoning benchmarks quietly assume you're working in English or a handful of major languages, and that gap matters if a Lakota-focused evaluation suite ever needs to test reasoning tasks rather than just translation or retrieval. And a paper on Mongolian in its traditional script deals with what happens when a language is digraphic, meaning it can be written in two different scripts, and how that script-induced ambiguity complicates translation. That's a nice, narrow analog for anything involving Lakota orthography variation across different writing systems and community conventions; the same class of problem, just a different language. Nothing this week is a must-read-immediately item, but the reasoning-in-English-first result and the RAG faithfulness evaluation are probably the two worth an actual skim, since both suggest concrete evaluation methodology rather than just a data point about some other language pair. Turning to language technology news, this week's roundup circles a single question: when a language moves onto a digital map, a phone app, or a speech recognition model, who actually holds the pen. Start in New Zealand, where Google Maps just shipped a new voice for pronouncing te reo Māori place names, after years of complaints that the app was mangling city and town names across the country. Google didn't build this alone. It partnered directly with Te Taura Whiri i te Reo Māori, the Māori Language Commission, the government-backed body charged with protecting the language. The new voice was trained on a Kiwi speaker over roughly six months, and it's rolling out over about two weeks, with street and road names promised in a later update. The Commission's chief executive, Ngahiwi Apanui-Barr, called it a step toward securing te reo Māori's future in the digital age. What's worth noting here isn't the pronunciation model itself — plenty of tech companies have shipped custom voices — it's the shape of the partnership. Google didn't scrape audio and train in-house without asking. It went to the nationally recognized language authority, worked within that body's oversight, and let that oversight shape both the training data and the rollout timeline. That's a sanctioned model of language technology: a big platform, a small language, and an institution empowered to speak for that language sitting at the table as a partner rather than an afterthought. Contrast that with a project out of the American Southwest, much smaller in scale but built entirely from the other direction. A Navajo man and former educator named George R. Joe has spent three years building an app called Tribal Trailz — a GPS-activated audio tour that plays as you drive historic routes like the stretch from Gallup to Flagstaff, or Flagstaff to Phoenix. As the app tracks your location, it narrates cultural and historical context about Navajo, Zuni, and Acoma lands and landmarks along the way, aiming directly at correcting the misconceptions tourists carry about Native peoples when they pass through this country. He's now in phase two, adding segments for Albuquerque and Santa Fe. There's no corporate partner here, no national language commission, no six-month training pipeline with an institutional sponsor. It's one person building infrastructure to control the narrative on his own homelands, on his own timeline. Put next to the Google Maps story, these two items frame the same underlying choice two different ways. One path to sovereignty runs through partnership with a platform that has resources but needs institutional legitimacy to use them well. The other runs through an individual or small group building outside the platforms entirely, slower, but with the terms fully in hand from the start. Neither path is obviously right for every case — a search giant's speech synthesis reach is hard to replicate solo, but going it alone means never negotiating away control in the first place. Zoom out from these two projects and you land in the analytical pieces published this week, which are trying to name what's actually going on at a systems level. One, published as Australia marked fifty years of NAIDOC Week — the country's annual week honoring Aboriginal and Torres Strait Islander history and culture — argues plainly that artificial intelligence risks becoming just the latest extractive force aimed at Indigenous knowledge unless it's built on consent, credit, and return from the outset. The piece points to something concrete happening right now as a counterexample done well: Aboriginal medical clinics in regional Western Australia are trialing AI-assisted screening for diabetic retinopathy, a diabetes complication that damages eyesight, and doing it in a way the author holds up as evidence that this can work when Indigenous Knowledge Systems and Indigenous data sovereignty are treated as mandatory design constraints rather than an ethics checkbox bolted on afterward. Notably, the specific data-sovereignty frameworks named there are the CARE Principles — collective benefit, authority to control, responsibility, and ethics — and OCAP, meaning ownership, control, access, and possession, a framework built by First Nations communities in Canada. Those aren't abstract ideals. They're specific, enforceable claims about who gets to decide what a model is trained on and what happens to the data afterward. A second piece, published just yesterday, tries to name the bigger contest these frameworks are fighting inside of. It frames the AI era as a three-way struggle over sovereignty itself — state sovereignty, corporate sovereignty, and what it calls Indigenous techno-sovereignty — and argues that older, nation-state-centered ideas of sovereignty don't capture how power actually moves through AI, because that power flows through corporate-controlled infrastructure that no government fully commands. As a working example of Indigenous techno-sovereignty already in practice, it points to Te Hiku Media, the Māori media organization whose speech recognition model for te reo Māori is now reported at ninety-two percent accuracy, built and governed entirely within Māori control rather than handed to an outside company to build on their behalf. It also cites the First Languages AI Reality project as a parallel effort. The piece's prescription is federated architectures, data trusts, and permanent Indigenous seats at the tables where AI standards actually get written — not consultation after the fact, but a standing seat before the standard exists. That's a meaningfully different claim than "partner well," which is what the Google Maps story shows. It's a claim about needing your own infrastructure and your own governing institution, one that doesn't depend on a partner's goodwill continuing past the current product cycle. Which is exactly why the article coming out of a UN side event on South Asia lands as the hard counterweight to everything above. Held on July sixth as part of the broader UN Global Dialogue on AI Governance, the session asked whether South Asian nations are building genuine sovereign AI capacity or something that looks like sovereignty but functions as a more sophisticated form of dependency. The discussion went directly at a problem that the Te Hiku Media and Māori Language Commission stories make look almost solved by comparison: Indigenous and tribal communities' ecological knowledge, territorial records, and oral traditions are entering AI systems across South Asia without consent, across borders, with no single actor anyone can hold accountable — this despite those communities' rights under the International Labour Organization's Convention one-sixty-nine on Indigenous and tribal peoples, and under the UN Declaration on the Rights of Indigenous Peoples. Read the two stories about Aotearoa alongside this one and the contrast is stark. In one country, a national language commission has enough institutional standing that a company like Google comes to it as a partner and builds around its terms. Across a different set of jurisdictions, oral traditions and ecological knowledge are simply flowing into training pipelines because no accountable authority exists to say no, or to say yes on defined terms. Same underlying technology, same underlying question of consent — completely different outcomes, and the difference is entirely about whether an institution with real authority exists and is recognized. Line all five of these up and a pattern falls out that's directly relevant to the kind of community-governed work this show follows for Lakota. The strongest position, in every one of these stories, isn't "we got invited to the AI project." It's "we already have the institution, the data, and the standard, and outside partners work within our terms or not at all." Te Taura Whiri wasn't consulted after Google trained a voice — it appears to have shaped the training from the start, over six months, with a defined rollout. Te Hiku Media didn't ask a tech company to build Māori speech recognition and hope for the best — it built the model itself and reports its own accuracy number. George Joe didn't wait for a museum or a state tourism board to greenlight a Native perspective on Route 66 history — he spent three years building the infrastructure himself. And the failure mode, the one South Asian tribal communities are living right now, isn't malice from any single bad actor — it's the plain absence of an institution empowered to consent or refuse on the community's behalf, which lets extraction happen simply because nothing is in place to stop it. None of that is a finished picture — announcement and demonstration are still two different things in most of these stories, worth holding apart. Google's voice rollout is genuinely shipping, in the next two weeks, which makes it one of the more concrete items here. The ninety-two percent figure for Te Hiku Media's speech recognition is a reported number, not something this roundup can independently verify from what's public this week. The Aboriginal diabetic retinopathy screening trials are described as trials, not deployed clinical tools. And the South Asia session was a discussion, a diagnosis of a governance gap, not yet a proposed fix with any institution attached to it. But taken together, the throughline is clear enough: sovereignty over language and knowledge data isn't a settled legal category waiting to be applied. It's built, case by case, the same way Te Hiku Media built its own model and the Māori Language Commission built the standing to shape Google's — through an institution that exists before the AI company shows up asking for training data, not one improvised after the fact. That's the roundup for this week. Language-tech work like this moves slowly and mostly out of public view, so it's worth pausing to notice when it adds up: better documentation, better tools, more control staying with the communities the languages belong to. Thanks for spending this time on it, Mason. Rest up, and I'll have another batch of research and language-tech news for you next week.