← Back to Resources

A Guide to Government AI for Agency Leaders

Hands typing on a laptop in front of glowing server racks in a dark room, cool and warm light accents

The AI mandate has arrived at your agency: an executive order, a federal action plan, and an OMB memo that expects an enterprise strategy for responsible AI use. The catch is that the data the models would touch does not leave your control. Public records are subject to FOIA, CUI carries marking and handling requirements, and classified material lives on networks where only a closed enclave is allowed. So every AI conversation in the agency runs on a single question: where does the data sit, and can you prove it?

This guide breaks down how agencies adopt AI under those constraints: the deployment ladder that data sensitivity forces, the workloads agencies actually deploy, the cost frame of TCO versus egress, and how to procure an AI platform that stays in-house, including the authorization artifacts to demand in an RFP.

Government AI is the use of artificial intelligence inside federal, state, local, and tribal agencies — where public records, controlled unclassified information, and classified data all remain under agency control. Unlike commercial deployments, the central constraint is not capability: it is sovereignty. Agency data, records subject to the Freedom of Information Act, and classified material must stay within boundaries the agency can prove, or the deployment does not happen at all.

The pressure to deploy is real and dated. Executive Order 14179, “Removing Barriers to American Leadership in Artificial Intelligence” (January 23, 2025), set the policy of “sustain[ing] and enhanc[ing] America’s global AI dominance” and directed development of an AI Action Plan within 180 days. That plan, “Winning the AI Race: America’s AI Action Plan,” was released by the White House on July 23, 2025, and identifies more than 90 federal policy actions. OMB Memorandum M-24-10, executed March 28, 2024, requires each CFO Act agency to develop an enterprise strategy for responsible AI use and to follow minimum practices when AI affects the rights or safety of the public. For agency leaders, the mandate is no longer whether to adopt AI — it is how to adopt it while keeping control of the data.

Why public-sector AI differs from commercial AI

Four structural differences shape every decision that follows.

Data classification, not just security. Federal information spans public records, Controlled Unclassified Information (CUI), and classified material. CUI is a defined program: established by Executive Order 13556 in November 2010 and implemented through 32 CFR Part 2002, with agency data reported annually. CUI carries marking and handling requirements that a vendor’s cloud region selection cannot satisfy on its own. Classified data sits above that, on networks such as SIPRNet, where only a closed, self-contained enclave is permitted.

FOIA-adjacent records obligations. Records created by an agency — including records produced with vendor assistance — can be subject to the Freedom of Information Act, 5 U.S.C. § 552, which entitles the public to access agency records absent an exemption. When a vendor holds agency-generated data, the agency must be able to retrieve it, produce it, and explain what happened to it. “The data lives in the vendor’s cloud” is not a defensible answer to a records officer, a FOIA request, or an audit.

Budget cycles, not annual renewals. The federal fiscal year runs from October 1 to September 30, and agency IT funding is largely discretionary, set by Congress annually. An AI deployment that costs $200,000 a year must be modeled as a multi-year appropriation request with justification, not as a card on file. Multi-year cost behavior, not the sticker price, is what the budget office evaluates.

Procurement is the architecture. Federal purchase decisions run through the Federal Acquisition Regulation and established vehicles. OMB Memoranda M-25-21 and M-25-22, referenced by GSA’s Buy AI program, direct agencies toward efficient, governed AI acquisition, and GSA now lists AI products and services as standard contracting options. In practice, this means the deployment model has to be specified, defensible, and vendor-verifiable before the RFP goes out — not negotiated after a demo impresses a stakeholder.

Deployment options by data sensitivity

The first architecture question is where the data may sit. The answer determines everything else — security baseline, certification path, cost model, and which vendors can even bid. The standard ladder looks like this:

Data sensitivityDeployment modelCertification / authorization path
Public / internal, low impactPublic cloud (IaaS/PaaS/SaaS), commercialFedRAMP Low authorization
CUI / mission data, moderate impactFedRAMP-certified public cloud, or government/private cloudFedRAMP Moderate; agencies may leverage an existing agency ATO for an AI service
High-impact unclassified CUI, DoD environmentsDoD-approved commercial government cloudDoD Cloud Computing SRG Impact Level 5 (IL-5)
Classified (up to Secret)On-prem or enclave data center inside the agency’s classified networkDoD Impact Level 6 (IL-6): a closed, self-contained enclave connected only to the classified network (e.g., SIPRNet); no commercial-cloud path
Highest sensitivity / disconnected requirementsAir-gapped on-prem deployment, no external network pathAgency-specific ATO with physical isolation; no external authorization body applies

Three consequences follow from this table. First, FedRAMP is a gate, not a nice-to-have. FedRAMP is a GSA-operated federal program for the security authorization of cloud services, and it categorizes cloud offerings at Low, Moderate, and High impact levels based on FIPS 199. GSA’s procurement guidance is explicit that all cloud service providers used by the federal government must be FedRAMP-authorized or in the process of obtaining authorization — and it recommends reusing an existing agency authority to operate rather than starting from scratch. Second, the ladder is not a menu. CUI rules mean some datasets simply cannot go to a commercial region; DoD impact levels mean IL-5 and IL-6 workloads run under different security requirements guides and different vendor ecosystems. Third, the top of the ladder is on-prem by construction. Classified and air-gapped environments are closed, self-contained, and disconnected — which is why the sovereign AI discussion (see what is sovereign AI) and on-premise AI matter to government buyers long before they matter to commercial ones. For defense-specific environments, sovereign AI for defense covers the IL-5/IL-6 authorization path in detail.

The workloads agencies actually deploy

Federal AI is not speculative. The workloads that dominate agency roadmaps cluster in four areas:

Document processing and records management. Agencies hold massive backlogs of records, correspondence, and case files. AI-assisted document triage, classification, summarization, and FOIA request support — locating responsive records, identifying exemptions, generating summaries — is the most common first deployment, because the data is internal, the workflow is repetitive, and the upside is measured in staff hours recovered. The records obligations cut both ways: the same documents the AI helps process are the documents the agency must be able to produce under FOIA.

Citizen services. Public-facing intake, routing, and response — permit applications, benefit inquiries, 311-style services, multilingual help desks. The sensitivity here is lower (often public or PII-level), which makes cloud deployment viable, but the service-level and accessibility expectations are higher than in commercial settings because the agency is the only provider.

Predictive maintenance of public assets. Bridges, water systems, rail, and fleet vehicles generate sensor and inspection data that predicts failure before it happens. This is a classic on-prem or edge-adjacent workload: the data is local, the model benefits from staying near the asset, and the deployment avoids egress costs that would make continuous telemetry uneconomic.

Situational awareness and operations. Fusion of feeds for emergency management, infrastructure monitoring, and threat picture — the workloads that push agencies into the higher rungs of the sensitivity table, where air-gapped and classified deployment models apply.

A pattern worth noting: nearly all four workloads are regulated AI workloads in the practical sense — the deployment is governed by a specific set of requirements (CUI handling, ATO, audit logging) that the vendor must demonstrate, not merely claim.

The cost frame of TCO versus cloud egress

The commercial AI price frame is per-token or per-seat SaaS billing. The government frame is different, in four ways:

  1. Multi-year budget line, not annual renewal. A FY-cycle deployment is scored on three to five years of total cost of ownership — license, hosting, integration, security review, and operations. The budget office will model year two and year three pricing explicitly; usage-based bills that “can grow quickly without proper monitoring” (GSA’s own words) are a disqualifying risk if the envelope is not modeled.
  2. Egress is a real line item. Predictive-maintenance and situational-awareness workloads move data continuously. Routing that telemetry to a remote cloud for inference and back produces recurring egress charges that can dominate the license cost — a structural argument for processing at the edge or on-prem, where the data never leaves the facility.
  3. Authorization has a price, and reusing it saves money. A FedRAMP authorization and an ATO consume assessor time, staff time, and calendar weeks. GSA explicitly recommends leveraging existing authorities where available. An in-house platform that already carries the agency’s ATO avoids re-running the authorization for every new model or workflow.
  4. Procurement overhead is front-loaded. Security review, contract negotiation, and vehicle compliance consume staff months before the first dollar of subscription. Vendors who come with FedRAMP status, an ATO package, and pilot-ready environments shorten this phase materially — which is a legitimate part of the price comparison, not a tiebreaker.

How to procure AI in government

The RFP is where sovereignty gets decided. The practical checklist:

  • Specify the deployment boundary in the statement of work. State explicitly where data is processed and stored: FedRAMP-authorized cloud, agency-operated data center, or air-gapped. “Vendor’s discretion” is not a requirement — it is a waiver of control.
  • Require the authorization artifacts, not just the claim. FedRAMP authorization letter (and impact level), current ATO or provisional ATO path, and for DoD workloads the DoD Cloud Computing SRG impact level (IL-5 or IL-6). Verify against the FedRAMP Marketplace, which lists every FedRAMP-certified service.
  • Demand data residency and portability language. The agency must be able to retrieve all data, models, and logs on contract end, and the vendor must not use agency data to train shared models without written authorization. Why enterprise data sovereignty matters more than ever frames the stakes of that requirement.
  • Require audit and records support. Immutable audit logs for model inputs and outputs, plus a contractual commitment to support records production (including FOIA responses) for agency-generated outputs. AI agents in regulated industries covers the compliance architecture side of that requirement.
  • Model the multi-year cost in the evaluation criteria. Weight three-year TCO, egress/egress-avoidance design, and usage-cap behavior explicitly in the scoring matrix rather than comparing sticker prices.
  • Structure a pilot-to-production path. GSA’s guidance recommends evaluating solutions in testbeds, sandboxes, or pilot programs with a small user group before large-scale purchase. Specify the pilot’s exit criteria (accuracy threshold, ATO status, cost confirmation) so the pilot is a decision, not an extension.
  • Engage the full governance bench before the RFP ships. GSA’s guidance names the coordination set: Chief Information Officer, Chief AI Officer, Chief Data Officer, Chief Information Security Officer, and Chief Privacy Officer. M-24-10 requires each CFO Act agency to develop an enterprise strategy for responsible AI use — the RFP should map to it.

Procurement is also where the market gap shows. Most incumbent AI vendors frame public-sector offerings around their cloud: the pitch is platform and models, and the deployment boundary is whatever the cloud allows. When an agency’s data cannot leave its own facility, that framing does not survive the requirements review. The sovereign framing — platform, model, and data all inside the agency’s control, with the authorization artifacts to prove it — is the frame that maps onto the actual sensitivity ladder. That is the core of the sovereign AI argument, and the one worth testing against a working system before an RFP is finalized. How to evaluate sovereign AI platforms walks through the criteria for that test. See the demo.

Last verified: 2026-09-05

Why Government Agency Leaders Choose Shakudo to power their on-shore AI programs

  1. Data sovereignty by design

    The AI stack runs inside the agency's own environment, so public records, CUI, and classified material never leave the government boundary.

  2. Runs in federal and agency infrastructure

    Shakudo deploys the platform inside agency-controlled facilities and networks, from agency data centers to air-gapped enclaves, so the deployment boundary is the agency's to define.

  3. Audit-ready orchestration

    Model inputs, outputs, and routing decisions are logged, so agency records officers and auditors can reconstruct what the system did with agency data.

  4. Vendor-agnostic model serving

    Model serving is composed from components the agency selects, not locked to a single vendor's stack, so the government keeps exit paths across contract cycles.

Talk to us

Frequently asked questions

Can the stack run inside agency infrastructure?

Yes. Shakudo deploys inside agency-controlled environments, from agency-operated data centers to air-gapped enclaves, so the deployment follows the agency's sensitivity ladder rather than a vendor's cloud offering.

Does agency data leave the government boundary?

No. In an on-shore deployment, models, prompts, outputs, and records stay inside the agency's own infrastructure. The boundary is defined by the agency and can be verified in the deployment model.

What is the audit and compliance posture?

The platform supports the audit and records requirements the guide describes, including immutable logs of model inputs and outputs, and runs inside the environments agencies need to obtain an ATO or a DoD SRG impact level.

How do agencies typically procure this?

Through the Federal Acquisition Regulation and established vehicles, including the AI products and services GSA lists as standard contracting options. The deployment boundary and the required authorization artifacts are specified in the statement of work before the RFP ships.

Don't miss these

Ready to put this into practice?

Get Started