Public Sector Advisory Without Vendor Bias

Chart illustrating why public sector advisory must be independent of the AI vendor being evaluated

GAO reviewed 13 AI acquisitions this spring at the Departments of Defense, Homeland Security, and Veterans Affairs, plus the General Services Administration. None of the four agencies had a system for capturing what they learned from those deals. Not one.

That finding sits inside a report most people outside federal acquisition offices will never open: “Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements.” The title undersells what it documents. Federal AI use more than doubled between 2023 and 2024, and the agencies buying it are still solving the same problems contract by contract, because nobody wrote down what went wrong the last time.

The examples in the report are specific enough to sting. The VA retired its SoKAT suicide-prevention tool in January 2023 after concluding it didn’t outperform the tools already in place, and never documented why, so the next program office evaluating a similar vendor starts from zero. FEMA couldn’t share certain geospatial model outputs with state partners because the agency hadn’t secured the data rights at the time of award. The vendor had also told FEMA the model struggled to tell one type of dwelling from another; that feedback never made it into a written record either. Army officials, meanwhile, discovered that proposed licensing fees for AI on the XM-30 program were running near $300,000 per vehicle, per year, a number nobody had built into the original business case.

None of that was information the government lacked. It was information sitting with the people who sold the system, whose contract renews if the program looks successful and gets harder to defend if it doesn’t.

A newer data point makes the same case from a different angle. SAS and IDC surveyed government organizations worldwide this year and found only 6% operating in what they call the “ideal state”: high internal confidence in their AI paired with AI that is demonstrably trustworthy. That’s the lowest share of any industry in the study, behind banking, insurance, and health care. Thirty-eight percent of government organizations are simultaneously overrelying on AI and underinvesting in the safeguards meant to catch it when it fails. IDC’s Chris Marshall put it plainly: agencies are “moving quickly from AI experimentation to operational use, but trust can’t be assumed.” Confidence, in other words, is outrunning verification, and almost nobody outside the vendor relationship is positioned to close that gap.

The same study found government running below the global average on its Trustworthy AI Index, at 15.3% compared with 19.8% worldwide. Every region cited the same root cause: no centralized, well-governed data foundation to build on. Governance is part of this picture, but it isn’t the whole of it. An agency can write a sound AI policy and still sign a contract that hands away its own data rights, because policy work and deal review are different disciplines, staffed differently, and rarely sit with the same person.

Why the same mistakes keep showing up

GAO grouped the recurring trouble into two buckets. Three are strategic: access to subject matter experts, protection of government data and intellectual property rights, and acquisition timelines built for hardware procurement, not for software that changes weekly. Three are programmatic: how requirements and contract terms get written, whether testing happens early enough to matter, and whether anyone priced the full cost of the system rather than just the initial fee.

Together, those six gaps describe an acquisition office negotiating against a counterparty with far more information about its own product. OMB’s own guidance already assumes that imbalance exists. It tells agencies to test proposed AI solutions before award and keep monitoring them afterward, using validation data the vendor cannot see. That instruction is a tell. If the rulebook requires evaluation data hidden from the seller, the government has already conceded that the seller’s own evaluation isn’t the one worth trusting.

What an independent advisor catches that a vendor won’t

A vendor’s pitch is built on data from other clients, tuned to perform well in the room. What happens on the agency’s own data, in the agency’s own environment, is a separate question, and it’s usually the one nobody in the sales cycle gets paid to ask. FEMA’s dwelling-classification problem is a plain example. It wasn’t concealed. It was a limitation the vendor disclosed once and then watched disappear from the record, because no one outside the deal had a reason to keep raising it.

An advisor with no stake in whether the contract closes has exactly one reason to raise it again: the program will fail quietly, months later, if someone doesn’t. That’s a different posture than the one built into most sales cycles, and it’s worth naming directly. Vendors are not doing anything improper when they emphasize what works. They are doing their job. The gap only closes when someone whose job is different sits at the same table.

Where independent review belongs in the acquisition timeline

The fix isn’t a single gate. It’s three, spaced across the life of the contract.

Before the RFI or RFP goes out, someone independent of the eventual vendor list should help define the data rights, intellectual property terms, and exit provisions the agency actually needs, not what a specific bidder is comfortable offering.

During testing, evaluation should run on data the vendor never touches, exactly as OMB already recommends, read by someone with no stake in a passing grade.

After award, someone should be responsible for writing down what worked and what didn’t, in a form the next program office can actually find. GSA’s own USAi effort shows this is solvable in practice: officials there wrote a privacy policy and contract language setting data ownership terms and limiting vendor access to chat interaction data before a single system went live. That took someone in the room whose only job was protecting the government’s interest.

What to ask before hiring an AI advisor for a government program

A short set of questions tends to separate an independent advisor from an extension of the sales process.

  • Does the firm also implement, resell, or take referral fees tied to the systems it evaluates?
  • Does it get paid the same amount whether the recommendation is yes or no?
  • Will it still be at the table during testing and post-award monitoring, or does the relationship end at contract signature?
  • Can it read the data-rights and intellectual property language in a draft contract, not just summarize a vendor’s capability deck?

An agency that can’t answer all four with confidence is buying an opinion that was never independent to begin with. That’s a harder standard than most procurement checklists apply, and it should be.

The four agencies in the GAO report were not short on technical talent. What none of them had, structurally, was a second, disinterested set of eyes on the deal before signature, someone paid the same whether the program worked or not. That is the specific gap our government and public sector advisory work is built to close. Agencies earlier in the process may find it useful to start with our government AI readiness assessment, and the evaluation questions we use to pressure-test vendor claims are the same ones laid out in our AI and XR due diligence checklist, built first for private equity deal teams and just as relevant to an acquisition office weighing its next AI contract.

Tags:

Comments are closed