Operation Epic Fury made clear that AI is now on the coronary heart of American warfighting. Central Command used Claude by means of Palantir’s Maven Good System platform to generate and prioritize roughly 1,000 targets within the first 24 hours of the marketing campaign, at an operational tempo that greater than doubled the opening section of the 2003 Iraq invasion. Over 38 days, the marketing campaign reached 13,000 complete strikes, based on Pentagon knowledge. Challenge Maven says that this system has accelerated focusing on even additional to five,000 in a single day. The query is whether or not the division can consider what it deploys earlier than it deploys it.
Over the previous 18 months, the Pentagon cleared away non-statutory obstacles to AI adoption that have been lengthy overdue for elimination. That work was mandatory, however inadequate. Constructing the division’s capability to facilitate broad adoption is the second half of the acceleration effort the Barrier Elimination Board began. The more durable job is constructing the coverage, workforce, and infrastructure capability for broad diffusion, since competitors activates how broadly a navy integrates a functionality, moderately than fielding essentially the most superior system first. With out that capability, adoption will keep confined to a handful of early-adopter workplaces.
The Pentagon’s AI Acceleration Technique, launched in January, drove a lot of that progress. It stood up a Barrier Elimination Board with authority to waive non-statutory necessities throughout testing, contracting, and hiring. It additionally required adoption of frontier fashions inside 30 days of launch, punting the how-to future steering. The technique is constructed on an April 2025 govt order streamlining acquisition and the July 2025 AI Motion Plan, making a political mandate for speedy AI adoption. Congress adopted with the FY26 Nationwide Protection Authorization Act.
The division set July because the deadline for preliminary demonstrations throughout its precedence AI deployments, or pace-setting initiatives. These demonstrations have been meant to show that the technique had unlocked actual functionality, as a substitute of merely accelerating the fielding of programs the division can’t but undertake at scale. They need to have introduced proof that the expertise was being tailored by means of speedy iteration cycles in near-operational settings and competitors between small groups overtaking the division’s centralized planning. Nonetheless, the division missed its personal deadline for preliminary demonstrations. Nor has it made any statements in regards to the standing of the initiatives.
Generative AI Requires New, Tailor-Made Techniques, Insurance policies, and Investments
Not like the slender AI programs the Pentagon has fielded for years, which do one job and keep fastened, generative fashions do many duties and alter consistently. That renders the standard “take a look at now, deploy perpetually” method out of date and requires new coaching investments to assist operators keep away from over-trusting or under-trusting AI programs.
Frontier mannequin reliability varies enormously by job. On the HalluLens benchmark, frontier fashions hallucinate on roughly 27 to 85 p.c of brief factual queries and invent solutions about nonexistent entities as much as 94 p.c of the time. Pentagon coverage has not stored up with this actuality.
Automation bias is well-documented in human-machine teaming analysis. As a result of these fashions generate articulate suggestions full with exact coordinates and prioritized targets, they threat inspiring extreme operator deference. Compressing analysis to fulfill a 30-day deployment mandate undercuts its personal function: a mannequin whose failure modes nobody has mapped is one operators be taught to not belief, and programs operators don’t belief get bypassed within the discipline.
On Could 1, following a dispute with Anthropic over the corporate’s makes an attempt to implement limits on using its Claude mannequin, the division expanded the roster of frontier AI distributors approved for the Pentagon’s highest-classification networks from one to eight. The Pentagon has not matched that growth with extra analysis infrastructure, workforce, or authorized capability.
Nationwide Safety Presidential Memo 11 Doubles Down on Pace Over Techniques
Nationwide Safety Presidential Memorandum 11, signed June 5, bolstered this “undertake now, construct the mandatory programs later” method. The memo known as for the elimination of “pointless obstacles to speedy deployment.” Assurance, or guaranteeing that AI applied sciences are “dependable, sturdy, steerable, and controllable” and observe relevant legal guidelines, insurance policies, and steering, was one of many memo’s 4 pillars. Nonetheless, it didn’t determine which workplace can be liable for this perform, nor did it direct funding for the work. As an alternative, it paired this emphasis on guaranteeing adopted AI is protected with intense operational stress by mandating the deployment of latest fashions in only a month. On this manner, it implicitly claimed to be advancing each adoption and assurance on the similar time, successfully pushing operational deployment far forward of unbiased testing.
There are two workplaces that might theoretically tackle the peace of mind function, and neither is in a position to take action. The Chief Digital and AI Workplace’s Accountable AI Workplace, the unit constructed to develop the division’s AI governance and assurance processes, has absorbed the job in follow, however lacks the capability: it started constructing clearer testing and analysis processes this spring, solely to lose employees to the deferred-resignation buyout program and the return-to-office mandate. The choice, the Director of Operational Take a look at and Analysis, has fared no higher: Secretary Pete Hegseth minimize its employees roughly in half in Could 2025 and gave it solely seven days to implement the change. The workplace has since dropped practically 100 packages from its oversight listing with out adapting its strategies to guage general-purpose fashions that replace each few months. The capability to independently consider generative AI doesn’t exist wherever within the Pentagon on the scale the mission requires.
The memo additionally mandated the termination of contracts with corporations that impose limitations on what the division can do with their expertise. The clause was a response to Anthropic’s insistence that it retain the power to restrict using its fashions for totally autonomous warfare or home surveillance, which resulted within the division declaring it to be a provide chain threat. Implicit in that place was a perception that neither the division, nor the broader U.S. authorities, might be trusted to make sure the expertise is used appropriately and lawfully. Whereas delegating such core coverage choices to the non-public sector is problematic in a number of methods, the inadequacy of the division’s personal assurance processes undermines the federal government’s place that it must be the only real arbiter of acceptable use of those instruments. Extra alarming nonetheless, the division has now written into coverage an incentive to pick out for distributors that impose fewer constraints on the very second it lacks the capability to independently confirm what any of them ship.
The memorandum ignored a fundamental level: operators undertake AI they belief and refuse what they don’t, and raised the danger that the navy authorizes programs for operational use earlier than anybody totally understands or secures them. The memorandum did direct the creation of standardized testing methodologies, a compute roadmap, and a reserve corps of out of doors technical consultants. These are the appropriate priorities. Correctly funded, they might pace adoption: a commander solely fields what has handed a take a look at they belief. However till Congress funds these directives, they continue to be aspirational.
The 5 Pillars of the Capability Downside
Frontier fashions replace each few months, and the division is just not staffed, geared up with ample compute, or methodologically ready to soak up them at that cadence. The brand new Barrier Elimination Board can’t waive its manner out of a capability downside.
There are 5 main shortfalls: workforce, compute, knowledge, analysis, and authority to function. The division doesn’t have machine studying engineers in ample numbers, acquisition officers who can write a consumption-based contract moderately than a firm-fixed-price, or operators skilled to catch a fluent-sounding mannequin’s fabrications. Even so, a safety clearance alone takes the higher a part of a yr earlier than an engineer can begin significant work. On compute, Pentagon-owned services run six-to-eighteen-month accreditation timelines, or greater than two years for categorized or complicated programs. On knowledge, the division lacks a unified structure that enables frontier fashions to entry the operational knowledge they require. The Pentagon usually can’t use knowledge generated by its personal operations to coach AI as a result of vendor contracts don’t safe these rights: the division funds the system and walks away with out proudly owning it.
On analysis, the division leans on vendor-provided benchmarks by means of a “rent-a-bench” mannequin as a substitute of proudly owning the methodologies and infrastructure that unbiased oversight requires. The latest publicly obtainable testing knowledge illustrates the stakes: As of 2024, Maven appropriately recognized objects at 60 p.c accuracy, in contrast with 84 p.c for human analysts in the identical 18th Airborne Corps evaluations. The division has no unbiased benchmarking infrastructure to measure whether or not that hole has narrowed, at the same time as intelligence officers have publicly dedicated to a aim of utilizing Maven to make “1,000 high-quality [targeting] choices” per hour. With two-year-old knowledge that predates Operation Epic Fury, nobody exterior this system can say whether or not the human-analyst hole has since closed, widened, or reversed.
On authority to function, the method of certifying a system safe sufficient for a authorities community, delay is inbuilt. A single authorization can take twelve to eighteen months, a cadence constructed for programs that change not often, not for fashions that replace each few weeks. The division has began to reply: Congress has written steady authorization and cross-service reciprocity into regulation. However reciprocity stays largely aspirational: a mid-2026 survey of division decision-makers discovered none reporting authorization timelines beneath six months, with “reciprocity” usually nonetheless that means a full redo of the previous overview.
All 5 gaps will worsen if the one instrument the Pentagon makes use of is waiving processes. The clearest threat is the navy’s incapability to confirm whether or not an AI answer does what the seller claims, or does it safely.
What Failure Seems to be Like
The division has seen what occurs when it fields automated programs whose failure modes nobody has mapped. In 2003, Patriot air protection batteries shot down a British Twister and a Navy F/A-18 over Iraq, killing three aircrew. The batteries operated in a largely computerized mode that conditioned operators to belief system outputs unconditionally. A Protection Science Board job power later discovered that when the system’s assumptions stopped holding, operators had no solution to query what its sensors instructed them. The Patriot was a mature, narrowly scoped system with many years of testing behind it, whereas the massive language fashions the division is now fielding on 30-day timelines are neither mature nor slender.
When machines generate suggestions quicker than people can correctly consider them, overview will get rushed till it stops functioning as a examine. In keeping with an investigation by +972 Journal, Israel’s Lavender system flagged some 37,000 individuals in Gaza as suspected militants, with an error fee officers estimated at 10 p.c, whereas human reviewers spent about 20 seconds deciding who to kill. The Israeli navy disputes points of that reporting, however the sample is acquainted in automation bias analysis: when a system produces hundreds of suggestions with fluent confidence, the human within the loop turns into a formality. (Some is likely to be pondering of the U.S. strike on an Iranian faculty that killed 156 individuals, together with 120 youngsters, however the proof to this point factors to not a discrete AI error however to broader failures within the U.S. focusing on system, together with outdated intelligence and disconnected databases.)
Departmental focusing on doctrine spells out what that scrutiny is meant to seem like. Joint Publication 3-60 requires optimistic identification of a goal, a collateral injury estimate weighed towards navy necessity, and authorized overview earlier than a commander approves a strike. That solely works if analysts have time to do it significantly.
Set these precedents towards Maven’s tempo throughout Operation Epic Fury: Nobody, inside or exterior this system, can say how lots of the system’s object identifications throughout that marketing campaign have been right. When failure comes, it can seem like a strike on a constructing the mannequin confidently misidentified, authorized by an operator skilled to maneuver quick, in a warfare the place no analysis file exists to determine whether or not the error was foreseeable. The division can’t say how usually that’s already taking place, and that silence, moderately than a two-year-old accuracy determine, is the capability hole in its starkest kind.
Congress has given the division the authority to begin closing these gaps. The FY26 Nationwide Protection Authorization Act approved three provisions: a testing sandbox, a cross-functional analysis workforce, and a governance subcommittee to assign clear accountability for AI oversight. Not one of the three has been established, regardless of deadlines for all three having come and gone.
What the Pentagon Ought to Do Instantly
None of this slows the pace-setting initiatives. It’s going to maintain them from stalling as soon as the overdue demonstrations lastly arrive. The Secretary ought to designate workforce, compute, knowledge, and analysis infrastructure as an official pace-setting venture beneath a single accountable chief drawn from the Accountable AI Workplace. Whereas assigning this mandate to an already depleted workplace poses speedy capability challenges, formal pace-setting standing gives the precise mechanism wanted to power senior management visibility and unlock devoted resourcing. A pace-setting designation issues as a result of it comes with standing entry to the Barrier Elimination Board and metrics that the Beneath Secretary for Analysis and Engineering personally critiques every month. To maneuver sources moderately than paper, the venture wants two issues: price range authority for its chief, and metrics that depend outputs moderately than milestones: engineers employed and cleared, compute delivered at every classification degree, and knowledge rights clauses executed in new contracts. Give the venture two years to point out outcomes. If it hasn’t reallocated funds towards these hires, that compute, and people knowledge rights by then, the Secretary ought to ask Congress for devoted funding as a substitute. Constructing capability deserves the identical administration consideration that barrier elimination has obtained.
The Pentagon should concurrently pivot towards provisional safety accreditations because the default for quickly updating AI fashions. Defaulting to provisional accreditation may seem like one other concession to hurry over security, however cybersecurity checks and functionality evaluations reply essentially completely different questions. Accreditation, the authority-to-operate course of, determines whether or not this system would introduce cybersecurity threat to the community that hosts it. Current assessments overview a static configuration, so for fashions that replace each few months, the reviewed model is commonly out of date earlier than the paperwork clears. The division’s personal Software program Quick Observe initiative already treats steady monitoring as a sounder safety posture for fast-changing software program than a point-in-time snapshot. Functionality analysis asks a special query: does the system do what the seller claims, beneath what situations, and with what failure modes? That’s the gate the division can’t at the moment employees, and it’s also the gate that can not be relaxed. Rushing up accreditation solely is sensible if functionality analysis stays a separate, unbiased examine that also applies irrespective of how briskly accreditation turns into. A mannequin that clears safety overview in weeks however has by no means demonstrated accuracy towards an analyst baseline has no enterprise in a focusing on cell.
Congress ought to fund unbiased testing infrastructure for navy AI, housed in universities and federally funded analysis facilities. Congress has already moved right here: Part 224 of the FY26 Nationwide Protection Authorization Act directed the Division of Battle to determine a Nationwide Safety and Protection AI Institute at a college. A mandate constructed round workforce improvement and foundational analysis, although, is just not a standing capability-evaluation perform.
The Pentagon at the moment depends on distributors to evaluate their very own merchandise, a battle of curiosity that leaves the federal government depending on instruments it neither owns nor controls. The onerous query is what these facilities would consider and towards what requirements, when the methodologies barely exist. NIST’s Middle for AI Requirements and Innovation has solely revealed draft steering on automated benchmarking of language fashions. Shifting ahead, unbiased testing ought to measure target-recognition accuracy towards human analyst baselines on categorized operational knowledge, because the 18th Airborne Corps did in 2024. They need to run adversarial exams for hallucination and immediate injection, repeat them with every new mannequin model, and run human-machine teaming experiments measuring operator scrutiny at operational tempo. Establishments like MIT Lincoln Laboratory and Carnegie Mellon’s Software program Engineering Institute already do categorized analysis work and will anchor this, and the sandbox and analysis workforce Congress already approved are pure autos. Congress ought to fund them, home a part of the work exterior the Pentagon, and require reporting on outcomes.
Operation Epic Fury confirmed that the Pentagon can use AI to prosecute a warfare at scale. The preliminary pace-setting initiatives, in the event that they materialize, will present whether or not they can maintain that momentum throughout extra programs. But neither effort ensures the navy can confirm system accuracy in fight or catch vital failures earlier than they occur. Addressing this hole is vital to driving broad adoption and guaranteeing that it lasts.
Jordan Kane spent 9 years advising Congress and the Departments of Protection and State on the Particular Inspector Basic for Afghanistan Reconstruction’s Classes Discovered Program. She is an AI fellow with the Horizon Institute for Public Service and a former Fall Fellow on the Middle for the Governance of AI.
Hamza Chaudhry is AI and Nationwide Safety Lead on the Way forward for Life Institute, the place his work focuses on AI security and nationwide safety. His work and commentary on these subjects have been featured on CNN, Fox, Reuters, the Wall Avenue Journal, Politico, Axios, the Bulletin of the Atomic Scientists, and in collaborations with consultants on the Departments of State and Homeland Safety.
This piece relies on a longer article about U.S. navy AI readiness revealed within the Fashionable Battle Institute’s journal.
Picture: Amn Alenne Mojica by way of DVIDS
