A blog about what it actually takes to bring artificial intelligence into state government, written by someone who’s spent years sitting across the table from the people doing it. Most coverage of this topic swings between breathless hype and easy cynicism about bureaucracy. Neither matches what’s really happening. State CIOs have made AI their top priority for the first time in over a decade, and they’re doing it while managing tight budgets, sensitive data, and decisions that affect real people’s benefits, healthcare, and livelihoods. This blog is for anyone who wants to understand that work honestly — the genuine barriers, the smart people navigating them, and the states quietly figuring out how to move from pilot to production without cutting corners on the public’s trust.

Too Many Good Choices: A Framework for Selecting an AI Platform in State Government

Open any state CIO’s inbox and you’ll find the same story: a steady stream of outreach from OpenAI, Anthropic, AWS, Microsoft, Google, Salesforce, and a growing list of smaller, sharper startups — each with a genuinely capable GenAI offering, each with public-sector case studies, each confident theirs is the right fit. None of them are wrong to make that case. This is one of the most competitive, fastest-moving markets in technology history, and the competition is producing real, meaningful improvements every few months.

Which is exactly the problem. When every option on the table is credible, “which AI platform should we use” stops being a technical question and becomes a decision-paralysis problem. A state CIO doesn’t have the luxury of running ten parallel pilots and picking a winner the way a fast-moving startup might — not with the approval timelines, staffing constraints, and data sensitivity we’ve already covered in this series. So how do you actually choose, when every vendor in the room has a legitimate claim on being excellent?

Before reading further — how is your organization deciding between AI vendors today? A formal evaluation matrix, a structured pilot process, or something more ad hoc? Most states I’ve talked to are still building that muscle.

Why the private-sector playbook doesn’t transfer

In the private sector, vendor selection is often a speed game: stand up a few proofs of concept, compare outputs, pick the best one, move on. That approach assumes you can afford to try and discard several vendors quickly. State government mostly can’t. Each new platform can mean a new security review, a new data-sharing agreement, a new procurement cycle, and new training for staff who are already stretched thin. Running that process five times in parallel isn’t cautious — it’s operationally impossible for most agencies.

That means the real question isn’t “which platform is best.” It’s “which path to a safely deployed system is shortest, given everything we already have in place.” That reframing changes the criteria completely — and it points toward a framework built around leverage, not just capability.

A six-step framework for evaluating AI platforms in a state government context.

Start with what you’ve already earned the right to use

Every state already has technology partnerships that took months or years to establish — cloud enterprise agreements, security authorizations, cooperative purchasing contracts. That existing relationship is worth more than it usually gets credit for in an AI evaluation.

Two mechanisms make this concrete. GovRAMP (formerly StateRAMP) runs on a “verify once, serve many” model: once a cloud provider clears the security review, any participating state can rely on that same authorization instead of re-running its own from scratch. NASPO ValuePoint, the multi-state purchasing cooperative, has already negotiated master agreements covering cloud and software with major providers — agreements any participating state can use without running its own competitive solicitation. If your state already has an active enterprise agreement and a completed security review with a cloud provider, adding an approved AI capability inside that same boundary is very often a matter of months, not the year-plus timeline a brand-new standalone vendor relationship requires. (Participation in both programs varies by state and agency, so it’s worth confirming your own state’s specific status rather than assuming coverage.)

Does your state already have an active enterprise agreement with a major cloud provider that could plausibly host an AI capability — and has anyone in your organization checked what’s already available inside it?

You don’t have to choose a vendor and a model at the same time

Here’s a distinction that gets lost in most vendor conversations: choosing a foundation model and choosing a new vendor relationship used to be the same decision. They no longer have to be.

The major cloud platforms states already work with increasingly offer multi-model gateways — a single, already-authorized environment that gives access to several leading foundation models side by side, under the same security boundary, identity and access management, and audit logging a state’s security team has already approved. That means a state can evaluate which model performs best for a specific task — summarizing casework notes, drafting citizen correspondence, flagging anomalies — without opening a brand-new procurement and security review for each model it wants to test. (The specific models and authorization levels available inside any given gateway shift often enough that it’s worth confirming current availability directly with your cloud provider rather than treating any published list as current.)

Trust the facts: prioritize what’s proven in production

Every platform will show you an impressive demo. Far fewer can point to a comparable deployment, at comparable scale, handling comparably sensitive data, in production today. That distinction should carry real weight in the framework — arguably more than any other single factor.

This is where Pennsylvania’s ChatGPT Enterprise pilot, discussed in our last post, is worth revisiting from a different angle: it wasn’t just a successful rollout — it’s now a reference point other states can point to when making their own case internally. It’s worth noting that Pennsylvania’s path (a direct, tightly scoped agreement with a single vendor) and the multi-model gateway approach described above are two different valid routes to the same goal, not the same path — the right one depends on what a state already has in place. What matters more than which route a state takes is the underlying question either way: “this configuration has already run at this scale, in this kind of environment, with these guardrails” is a far stronger argument to a skeptical security office than “the vendor says this will work.” When evaluating any platform, ask directly: where else, in a comparable government context, is this exact configuration already running — not a similar product, not a pilot, but this configuration, in production, today?

When a vendor cites a reference deployment, does your evaluation process actually verify it — call the reference, ask about the rough edges — or does the reference mostly go unchecked?

Verify the data-handling commitments

We covered this in depth in our last post, but it belongs in the framework explicitly: any platform under consideration should make a commitment that submitted data is not used to train or fine-tune models for other customers. That commitment needs to show up in the technical documentation. The clarity of that answer, and how easily a vendor can produce it in writing, is itself a useful signal about how seriously they take public-sector requirements.

Check the fit with what you already run

A platform that’s technically excellent but requires a brand-new, standalone integration is a heavier lift than one that plugs into systems already in place. If an agency already runs its casework or citizen-service operations through a platform like Salesforce, an AI capability embedded in that same ecosystem may involve far less integration risk than a separate, general-purpose AI tool bolted on from outside. Neither approach is inherently better — but “how much new integration work does this actually require” deserves to be asked before “how impressive is the demo.”

Preserve the ability to change your mind later

This space is still moving fast enough that the best model for a given task today may not be the best one in eighteen months. A framework built entirely around today’s leading model risks locking a state into yesterday’s choice. This is another reason the multi-model gateway approach is worth weighing heavily: an architecture that lets a state swap the underlying model without redoing the entire integration preserves flexibility that a single-vendor, deeply bespoke build does not. Optionality has real value in a market that’s still this early.

None of this is a knock on any single platform

It’s worth saying plainly: OpenAI, Anthropic, AWS, Microsoft, Google, Salesforce, and the rest are all building genuinely strong, differentiated public-sector capabilities, and that competition is good for every state watching the market. The point of this framework isn’t to declare a winner. It’s to argue that for government specifically, “which is the most capable model” is rarely the deciding question — nearly every serious platform is capable enough for the vast majority of state use cases. The deciding question is which path gets you to a safely deployed system fastest, with the least new risk, building on relationships and authorizations the state has already earned.

The framework, in short:

  • Start with what’s already authorized — existing GovRAMP authorizations and cooperative contracts shorten the path dramatically.
  • Separate the model decision from the vendor decision — multi-model gateways let you choose the best model without a new vendor relationship each time.
  • Require proof, not promises — prioritize configurations with referenceable production use over demo-stage claims.
  • Get data-handling guarantees in writing — in the contract, not the pitch deck.
  • Check ecosystem fit — integration cost is a real cost, not an afterthought.
  • Preserve the exit — favor architectures that let you change models later without rebuilding everything.