Independent Government AI Readiness Review

Chart illustrating the government AI readiness assessment gap between widespread AI piloting and the small number of states with real outcome evaluation in place.

Eighty-eight percent of federal decisionmakers call artificial intelligence a critical tool for modernizing their agencies. That figure comes from an Ernst & Young survey published in April 2026, and virtually nobody in government disputes it. Belief in AI is not the problem.

Here is the harder number from that same survey: half of all federal AI initiatives are still stuck in pilot or planning stage. Leadership conviction isn’t the bottleneck. Readiness is. And readiness turns out to be a narrower, less glamorous thing than most strategy decks suggest.

The pattern isn’t confined to Washington. Code for America’s 2026 Government AI Landscape Assessment tracked state governments across four stages: readiness, piloting, implementation, and impact. Nearly every state has now launched some form of AI pilot. Only seven have built the evaluation mechanisms needed to determine whether those pilots deliver measurable public value. Piloting is cheap. Proving impact is not, and most governments never get that far.

What Does “AI Readiness” Actually Mean for a Government Agency?

The confusion starts with the word itself. A pilot program is not evidence of readiness. It’s evidence of curiosity, funded at a scale small enough that nobody has to answer for what happens if it fails. Real readiness has to answer four separate questions before a single production dollar gets committed, and most agency self-assessments never get past the first one.

Who owns the decision when the model is wrong? Most agencies can name a project sponsor. Far fewer can name who owns model risk, drift monitoring, or the escalation path when an AI system produces a bad output in front of a constituent, an inspector general, or a congressional oversight committee.

Does the workforce exist to run this, not just approve it? The EY survey put a number on the gap: 44% of surveyed federal leaders named the skilled-labor shortage as their top modernization barrier, even as 95% said they’re actively investing in upskilling current staff. Training slides don’t close a staffing gap on their own timeline, and the two figures sitting next to each other say something the survey doesn’t spell out directly. Everyone agrees on the problem. Almost nobody has solved it yet.

Can the acquisition pathway actually deliver what got approved? GAO’s review of thirteen AI acquisitions across the Department of Defense, the Department of Homeland Security, GSA, and the Department of Veterans Affairs found agencies struggling to access technical experts capable of evaluating contractor proposals, and struggling just as much to understand the AI-related costs buried inside those proposals.

Is there a way to measure whether it worked? This is the question almost everyone skips, and it’s the one Code for America flagged hardest. Only seven states have built real evaluation infrastructure. Everyone else is renewing contracts on faith and calendar dates.

Why Did OMB’s 2025 Mandate Change the Conversation but Not the Outcome?

In April 2025, the White House and the Office of Management and Budget directed every federal agency to assess its AI maturity and build a plan to remove barriers to responsible use, with a 180-day clock attached to the requirement. That deadline did what deadlines do inside government: it produced a large volume of completed assessments. It did not, on its own, produce readiness.

This distinction is the one most compliance-driven exercises miss, and it’s worth stating plainly. An OMB-mandated maturity assessment is a paperwork requirement with a due date attached. An operational readiness assessment is a judgment call about whether an agency’s people, data, governance, and acquisition posture can actually support what leadership wants to deploy. An agency can complete the first exercise in full and still be nowhere close to the second. GSA’s response to this same gap — USAi, a shared government-wide platform for testing and scaling AI capabilities — is a genuinely useful tool. It is not a substitute for an honest look at whether a given agency’s people and processes are ready to use it.

What Do the States Getting This Right Actually Have in Common?

Code for America names seven states leading the field on 2026 maturity: Maryland, New Jersey, North Carolina, Pennsylvania, Texas, Utah, and Vermont. What they share isn’t budget size, and it isn’t the number of pilots each has running. It’s structure. Cross-agency governance that sits above any single department. Sandbox environments for controlled experimentation before anything touches a live system. And, most importantly, a built-in mechanism to measure whether a deployment actually worked before scaling it further.

That last piece is the one most readiness conversations still treat as optional. It shouldn’t be. An agency that cannot measure outcomes cannot defend its next budget request with anything more persuasive than enthusiasm, and enthusiasm is not a line item an appropriations committee funds twice.

What Should an Independent Government AI Readiness Assessment Actually Cover?

A readiness assessment run by a vendor tends to arrive at a familiar conclusion: the agency needs more of whatever that vendor happens to sell. An independent assessment starts from a different place, because there’s nothing riding on the answer except the answer itself. The version The Lion’s View runs for government clients — built on the same underlying logic as the organizational readiness framework applied with commercial and healthcare clients — covers five areas, examined in this order.

  • Governance ownership: who is accountable for model risk, and whether that accountability reaches agency leadership or stops at the engineering team.
  • Workforce capacity: whether the skills gap is being closed through hiring, partnership, or upskilling, and on what actual timeline rather than an aspirational one.
  • Data and technical infrastructure: whether the data feeding a system is documented, rights-cleared, and durable against a vendor’s pricing or policy changes.
  • Acquisition pathway: whether the agency’s contracting approach can deliver the capability that got funded, on the terms that were actually negotiated.
  • Outcome measurement: whether there is a defined way to know, six months after deployment, whether the thing worked and for whom.

Defense organizations face a compressed version of this same test, with a shorter runway and higher stakes. The Department’s AI-first mandate, examined in our review of the DoD’s readiness posture, forces the same five-part judgment onto a timeline that doesn’t leave room for a second attempt.

Why Does This Matter Before the Next Budget Cycle?

OMB’s maturity assessments aren’t a one-time exercise. They feed directly into how agencies justify technology spending in the next budget cycle, and the agencies walking into that conversation with a real, independently verified readiness baseline will carry more weight than the ones holding a completed checklist. The difference between the two won’t show up until someone asks a follow-up question the checklist was never built to answer.

That is the conversation The Lion’s View has with government and public-sector leaders through its government and public-sector advisory practice, before the RFP goes out and before the pilot has already stalled.

Sources: Ernst & Young federal AI survey (via Nextgov/FCW, April 2026); Code for America, 2026 Government AI Landscape Assessment; U.S. GAO, GAO-26-107859, Artificial Intelligence Acquisitions.

CATEGORIES:

No category

Tags:

Comments are closed