Every state has an AI pilot running somewhere right now. A chatbot answering resident questions. A tool summarizing casework notes. A model flagging anomalies in claims data. Ask around, and you’ll find dozens of these projects — quietly running in a sandbox, presented at a conference, maybe even written up in a press release.
Ask a different question — how many of those pilots became part of how the agency actually operates? — and the number drops fast. Recent research backs this up: nearly every state has piloted AI in some form, but adoption stalls well short of scaled, trusted deployment, and progress across states remains fragmented and inconsistent.
Having spent the last several years sitting across the table from state CIOs, CISOs, and program directors while they wrestle with exactly this problem, I want to use this first post to name the barriers plainly. Not the generic “government moves slowly” narrative — the specific, structural reasons AI initiatives get stuck. If you’re a state CIO, none of this will surprise you. But it’s worth saying out loud, because most of the public conversation about “government AI adoption” is written by people who have never sat in one of these approval meetings.
It’s also worth saying because the barriers below aren’t a story about neglect. AI just displaced cybersecurity — a priority that had held the top spot for 12 straight years — to become the #1 issue on state CIOs’ agendas for 2026. This isn’t an area getting ignored. It’s an area getting the most scrutiny, from people who are also being asked to do it with a shrinking budget.

Source: NASCIO, “State CIO Top Ten Policy and Technology Priorities for 2026”
Before we go further — where does your own state or agency sit today? Still evaluating pilots, or already running something in production? Keep that in mind as you read; I’d guess most readers land somewhere in the middle.
The approval process wasn’t built for this
Most state procurement and security review processes were designed for a world of static software: you buy a system, it does a defined thing, you certify it once, and it stays that way for years. AI tools don’t behave like that. The model changes. The vendor pushes updates. The way it responds to a prompt today may not be exactly how it responds next month.
That mismatch shows up in the numbers. Nearly every state has run some kind of AI pilot, but few have built the evaluation mechanisms needed to judge whether those pilots actually delivered public value — which means even successful pilots often have no formal on-ramp into production. Part of the reason is a step most people outside government have never heard of: the authorization to operate (ATO), the formal security sign-off a system needs before it can touch real data. ATOs were built for software that doesn’t change once it’s approved. A technology that could be evaluated in weeks in the private sector can take a year or more to clear that process — and by the time it clears, the underlying model has already moved on. CIOs aren’t slow-walking this because they’re risk-averse for its own sake. They’re applying a review framework built for stability to a technology whose defining feature is that it doesn’t stay still.
If you work in or with state government: how long did your last IT security review actually take, start to finish — and how much of that time was the technology, versus the paperwork?
Nobody has had time to build fluency in a technology that changes monthly
State IT and security teams are genuinely excellent at what they’ve always had to be excellent at: network security, identity management, legacy system integration, compliance documentation. What almost none of them have had is the time, headcount, or training budget to build deep, hands-on fluency with how large language models actually work.
Two terms are worth pausing on here, because the confusion between them drives a lot of the caution: training is the process of building or adjusting a model using large volumes of data, which happens before you ever interact with it. Inference is what happens when you send the model a prompt and it generates a response — no permanent change to the model occurs. A security reviewer who isn’t confident about that distinction has good reason to assume the worst-case version of both.
That’s not a skills deficit unique to government. It’s what happens anywhere a technology moves faster than an organization’s ability to formally train people on it — and the data backs this up starkly. Roughly 60% of state and local employees who are already using AI in their jobs report they’ve had no formal training on it at all.

Source: MissionSquare Research Institute, cited in “9 Key Risks of AI Adoption in State & Local Government” (2026)
When the people closest to a tool haven’t had structured time to build confidence with it, the default answer to “can we deploy this” becomes no — or “only in a form so locked-down it barely resembles the tool anymore.” That’s not an indictment of anyone’s competence. It’s a predictable, rational response to a training gap that budget cycles haven’t caught up to yet — and it’s compounded by the same fiscal squeeze showing up as the #3 CIO priority nationally. Asking teams to move faster on AI while also asking them to manage budget reductions is, for a lot of states, asking for two hard things with one hand tied behind the other.
Pennsylvania offers a useful counter-example of what closing that gap looks like in practice. Rather than opening AI tools to every employee at once, the state ran a deliberately narrow, governance-first pilot: 175 employees across 14 agencies, paired with structured training and hands-on support, before anyone else got access. A year later, that careful scope paid off — most participants reported a strongly positive experience and meaningful time savings on everyday tasks, and the state expanded access to more than 3,000 employees across 35 agencies. The lesson isn’t “move fast.” It’s that a small, well-supported pilot can build exactly the organizational fluency that a rushed, wide-open rollout can’t.
Would a narrow, well-supported pilot like Pennsylvania’s work in your organization — or does the culture push toward “roll it out to everyone” before anyone’s had a chance to learn it?
“We need our own environment” is non-negotiable — and it should be
Almost every state team I’ve worked with has said some version of the same thing: we are not putting resident data into a shared, multi-tenant model environment. (“Multi-tenant” simply means multiple customers’ data and workloads running on the same shared infrastructure — the default setup for most commercial cloud software.) States want isolated infrastructure, dedicated compute, clear data boundaries — something they can point to and say, definitively, “our data does not leave this environment.”
This is one of the more reasonable asks in the entire conversation, and vendors who show up expecting states to accept a shared SaaS model the way a private company would are going to keep losing these deals. State data includes benefits eligibility records, criminal justice information, health data, tax records — categories with their own statutory protections layered on top of general data security concerns. Isolated environments aren’t red tape for its own sake. They’re the only architecture that matches the actual risk profile of the data involved.
There are early signs the infrastructure to support this demand is catching up: the chief privacy officer role, once rare in state government, now exists (or has an equivalent) in 31 states, and many of those officers now sit at the center of AI governance and technology procurement decisions rather than being consulted after the fact. That’s a sign states are building the internal capacity to negotiate these environments confidently, rather than defaulting to “no.”
Does your agency have someone whose job is specifically AI governance or data privacy — or is that responsibility still spread thin across a general counsel’s office or an already-stretched CISO?
“Don’t train on our data” is about control, not paranoia
Closely related: state agencies are consistently firm that their data should never be used to train or fine-tune a model that other customers — or the public — might eventually interact with. (“Fine-tuning” means further adjusting a model’s underlying behavior using a specific set of data — different from simply asking it a question.) This isn’t a misunderstanding of how the technology works. It’s an accurate read of what’s at stake if it went wrong: a model that has absorbed patterns from confidential case data, surfacing in a way no one can fully predict or audit, in a context far outside the agency’s control.
Vendors that can clearly separate “using your data to answer your question right now” from “using your data to improve a model for everyone” earn trust quickly. Vendors that can’t — or that bury the distinction in a vague privacy policy — don’t get past this conversation. Pennsylvania’s pilot is instructive here too: the enterprise-tier agreement it used explicitly excluded employee submissions from any model training, and that guarantee was made explicit to participants up front, not buried in fine print discovered later. That clarity is very likely part of why the pilot earned enough trust to scale statewide.
Has a vendor ever been able to explain that distinction to your team clearly enough that it actually changed a decision — or has it always stayed too vague to act on?
The real fear isn’t the technology. It’s what happens to a person because of it.
This is the one that matters most, and it’s the one outside conversations tend to miss entirely. State agencies aren’t just processing generic business data. They’re making decisions that determine whether someone keeps their healthcare, their food assistance, their disability benefits, their housing voucher. When an AI system touches that decision — even indirectly, even just by summarizing a case file for a human reviewer — the question isn’t “will this save time.” It’s “what happens to a real person if this is wrong, and can we explain why it was wrong.”
That’s a fundamentally different bar than almost any private-sector AI use case. A retailer whose recommendation engine has a bad day loses a little revenue. A state agency whose AI-assisted process contributes to a wrongful benefits denial has caused real, sometimes irreversible harm to someone who had no ability to opt out of the system in the first place. Every hesitation I’ve seen from program staff — every insistence on a human in the loop (meaning a person reviews or approves the AI’s output before it affects anyone) — traces back to that asymmetry. It is, frankly, the right instinct.
Utah’s approach shows one way to honor that instinct without freezing entirely. Its Office of Artificial Intelligence Policy runs a formal regulatory “Learning Lab” — a sandbox where new AI use cases are tested under close supervision before they’re allowed to touch real decisions. Its first case study was AI in mental health care, one of the highest-stakes use cases imaginable, and the pilot was structured so a licensed physician remained in the loop for every exception, with the state actively monitoring safety outcomes throughout. It’s a template for how a state can say “yes, but carefully” instead of defaulting to “no” — testing the highest-stakes use cases under the tightest supervision, rather than avoiding them altogether.
If your state built a sandbox like Utah’s, what’s the first high-stakes use case you’d actually want to test in it — and who, specifically, would need to stay in the loop?
None of this is the technology failing. It’s the process catching up.
Here’s the thing that gets lost in most coverage of this topic: these aren’t excuses. They’re legitimate, well-founded concerns from people who understand their obligations to the public better than any outside vendor or consultant does. The states that will move fastest from here aren’t the ones that convince their IT and compliance teams to loosen up. They’re the ones that build review processes, isolated architectures, and governance models that actually match how this technology works — so that “yes” becomes a possible answer to a well-designed proposal, instead of “no” being the only safe answer to an ambiguous one.
The barriers, in short:
- Approval processes built for static software — ATO reviews designed for systems that don’t change, applied to models that update constantly.
- A training and fluency gap — roughly 60% of employees using AI have had no formal training on it, which pushes teams toward caution by default.
- The demand for isolated environments — a reasonable response to the sensitivity of the data involved, not red tape for its own sake.
- Refusal to let vendors train on agency data — a control issue, not a misunderstanding of the technology.
- Fear of impact on real people — the recognition that a wrong AI-assisted decision about benefits or eligibility causes irreversible harm, unlike most private-sector use cases.
That’s what this blog is going to dig into, post by post: not “AI will transform government” hype, and not “government is too slow” cynicism — but the real, specific mechanics of what it takes to get from pilot to production when the data is sensitive, the stakes are high, and the people affected by a bad outcome had no say in whether the system was used on them in the first place.
Next up: what a state-ready AI governance framework actually looks like — and why “more restrictions” isn’t the same thing as “better governance.”
Leave a comment