AI is making the ultimate goal in language access look possible: every resident is served in their own language, on time and under budget. Your team gets translations in minutes instead of weeks, and chatbots or AI interpreters answer residents in the non-majority languages you struggle to consistently staff.
Yet there are quite a few reasons to move carefully. AI is not infallible, and it takes a trained eye to catch translation mistakes. Case in point: during the 2021 vaccine rollout, Virginia’s health department ran its Spanish guidance through Google Translate. ‘The vaccine is not required’ came out, in Spanish, as ‘The vaccine is not necessary.” The English version explained that you weren’t legally required to get the vaccine. The Spanish-language version implies there was no reason to do so.
Many agencies are adopting these tools faster than they are writing rules for them, which increases the risk.
We’re here to tell you: You do not have to choose between speed and the safeguards. You just need to know which communications AI can handle today and which need a qualified human. Here, we take a look at the risks of ad hoc AI use in government language access programs and give you a framework for avoiding them.
Why Language Access Work Needs Its Own AI Rules
When you serve the public, your words, whether written or spoken, come with power, responsibility, and authority. Here are the main risks that keep AI from being a silver bullet for government agencies.
Accuracy Varies by Language
AI engines don’t perform equally in all languages. In a 2025 study, researchers compared AI translations of hospital discharge instructions against professional human translations. The AI matched the professionals’ error rate in Spanish, but it produced clinically significant errors in 92% of Somali sections.
When AI makes mistakes, the output often sounds accurate even when information is missing or altered. And performance can differ between dialects of the same language, too.
Effective Quality Control Requires Qualified Native-Speaker Linguists
An employee who reads only English has no way of evaluating the Spanish version of a page. And you can’t completely trust the AI to review itself. Automated Quality Estimation (AQE) uses AI to estimate the quality of translated content, and it’s a helpful tool in many situations. But it’s absolutely not a replacement for human review when the stakes are high. In one experiment, AQE missed two of nine medical translations that bilingual experts judged capable of causing harm.
Legal terms and culturally specific phrases add another risk. A translation can get every word technically right, and still give the reader the wrong meaning. Reliable review comes down to a qualified native-speaker linguist.
Data Protection Rules Still Apply
Anything an employee types into a consumer AI tool enters the public domain, including case details and personal information. These details can then be stored or used to train AI. And they can be leaked or unintentionally shared. In the most recent example of this, Google indexed a trove of “share this chat” URLs from conversations with Anthropic’s chatbot Claude. The chats, which contained personal details, then appeared in Google Search Results. Government work requires much stronger data protection than this.
Some Risks Appear Only After Publication
When an uncaught error in translation causes harm, many agencies have no one assigned to correct the error and answer for it. And each failure weakens residents’ trust in whatever the agency publishes next.
Where AI Is Already Being Used Successfully
The federal government’s 2025 AI inventory names more than 30 translation and interpretation systems across at least 10 agencies. Beyond the federal inventory, agencies at every level publish machine-translated information pages, draft multilingual outreach with generative tools, and translate routine correspondence automatically. At public meetings, the same systems produce live captions and speech translation. Behind the scenes, AI also routes and schedules interpreters to reduce the wait for the right professional.
Deployed thoughtfully, these systems can make a positive difference for constituents. When New Jersey added AI translation assistants to its unemployment insurance content, the time a Spanish speaker needed to finish an application fell from more than three hours to 28 minutes. Speed, scale, and availability are exactly what stretched language access programs are short on.
Not Every Use Case Carries the Same Level of Risk
The key is to use AI at the right times and the right places to expand access, while still tapping into human expertise when needed. We call this risk-based governance. The dividing line between using AI and not is what a mistake would cost the people you serve.
Lower-Risk Uses
The lowest tier is internal work: brainstorming, first drafts headed for professional review, terminology lists, summaries of non-sensitive documents. Mistakes here have little to no impact outside your agency.
Moderate-Risk Uses
Public webpages, answers to common questions, outreach materials, and appointment reminders make up the middle tier. AI-assisted workflows fit here when the agency runs per-language testing, human review, and a clear correction path.
High-Risk Uses
The top tier is anything that impacts rights, safety, or money. That includes benefits and eligibility decisions, emergency instructions, court and law enforcement interactions, immigration cases, public health directives, and notices about consent, appeals, or enforcement. Qualified human translators and interpreters carry this work, with AI in support at most.
Regulators are already drawing the same lines. Federal health rules require qualified human review of machine-translated critical documents, and the District of Columbia limits AI interpretation to low-stakes interactions.
What Responsible AI Governance Should Look Like
Rules organized by risk tier concentrate review where mistakes cost the most and leave internal, low-stakes work free to run. Agencies that write governance documents get more use out of AI, at lower risk, than the ones that improvise.
How to Build Your Own Risk-based Model
Use this checklist to get started:
- Inventory the AI already in use, including tools being used informally
- Classify your multilingual communications by risk tier, and describe how and when AI can be used for each
- Keep sensitive constituent data out of unvetted tools
- Name the approved, restricted, and prohibited uses for each tool
- Set performance and quality requirements for every tool
- Review vendor security, data retention, and model-training practices
- Test every tool in the languages and channels it will actually serve
- Define where human review is mandatory, in writing
- Record how reviews happen and who signs off
- Include language access professionals in AI procurement decisions
- Reassess the rules as tools and guidance change
- Build an escalation route from every automated channel to a qualified interpreter or translator
The Continued Role of Human Expertise
The risk-based model that we use at BIG shows your team when it’s beneficial to use AI and when they need human assistance. And they will need it, because AI is not a replacement for a highly skilled linguist. In this model, qualified translators and interpreters review and correct AI output, test tools before launch, monitor quality afterward, manage specialized government terminology, and take the handoff when an automated conversation stalls.
Operational Governance Must Keep Pace with AI Adoption
According to the Alliance for Innovation, 75% of agencies report AI use in some form, and only 14% have sufficient governance in place. The open question is how AI fits inside the language access requirements that federal and state law already set, alongside standing privacy, accessibility, and security obligations. Early guidance gives conflicting answers. California accepts machine translation of certain public materials as legally sufficient, while Oregon bars its agencies from AI translation at public meetings. Meanwhile, when AI use is not expressly forbidden, staff are adopting it one by one, often informally and without safeguards.
In government, the need for guidelines is urgent. And when guidelines do exist, AI policies often aren’t drafted around the unique requirements of language work: which languages to test, when a human reviews output, and how to protect sensitive constituent information.
Forward-thinking agencies can safely use AI to better serve constituents by creating their own internal governance now instead of waiting for guidance.
How BIG Can Help
BIG Language Solutions staffs the human half of that model for government agencies: qualified linguists who test, review, and correct AI output in more than 300 languages, and court-certified, healthcare-trained interpreters for the conversations that need one.
We also build AI solutions explicitly designed for the requirements of state, local, and federal agencies, including secure AI translation with human review when needed, and AI interpreting for simple over-the-phone queries.
LanguageVault®, our secure translation management platform, runs machine translation and human review inside one workflow, with the data retention and access controls your policy can name as approved. And for agencies building the policy itself, we consult on language access planning and compliance.
Contact us to build a language access program that pairs AI speed with qualified human review, so you can trust every word.



