2 October 2026

The Pentagon Has a New Test for Top Officers. It’s Not Clear Why.

The Bulwark | Mark Hertling

U.S. Secretary of Defense Pete Hegseth unveiled the Joint Warfighting Evaluation to assess colonels and captains for flag rank using written testing and artificial intelligence-assisted wargaming. Designed by former Marine Lieutenant Colonel Stuart Scheller, the initiative seeks to screen prospective general and flag officers outside traditional service promotion boards.

The Pentagon justifies the reform by invoking General George C. Marshall’s 1941 Louisiana Maneuvers as historical precedent for identifying operational commanders under battlefield stress. However, those historic maneuvers functioned primarily as large-scale combined-arms force experiments rather than isolated leadership exams. Modern combat training facilities like the National Training Center already evaluate operational command under realistic conditions. Furthermore, statutory joint qualification requirements established under the Goldwater-Nichols Act mandate career-long operational development across services. Standardized testing cannot simulate human command dynamics. Consequently, the Department of Defense must rigorously validate this wargaming instrument against experienced flag officers before institutionalizing new promotion metrics.

Comment

Synthetic wargaming environments isolate decision-making into discrete variables, failing to replicate the inter-service friction inherent to operational command under the Goldwater-Nichols Act. Strategic leadership within joint task forces relies on resolving inter-agency friction and managing delegated command authority rather than optimising tactical choices inside an algorithmic scenario. Evaluating senior commanders through isolated simulations treats theatre-level command as a closed technical system rather than an exercise in organisational leadership.

This technical reductionism misinterprets how command authority functions at instrumented centres like the National Training Center at Fort Irwin. At Fort Irwin, brigade combat team commanders are evaluated not on static option selection, but on their ability to sustain command resilience against free-thinking opposing forces under severe physical friction. Abstracting these dynamics into artificial intelligence algorithms risks prioritising tactical optimisation over the human judgment required for complex joint operations across combatant commands like USINDOPACOM.

Strategic Question for Discussion
If synthetic wargaming metrics displace live-force evaluations at the National Training Center at Fort Irwin, does simulated decision-making accurately reflect a commander's ability to handle operational friction under the Goldwater-Nichols Act, or does it privilege algorithmic problem-solving over human command resilience?
The pattern suggests that simulated environments fail to capture the friction, ambiguity, and human leadership demands inherent to field command. While synthetic algorithms efficiently test discrete cognitive choices, they cannot replicate the organizational friction experienced during force-on-force maneuvers at Fort Irwin. Consequently, relying on synthetic metrics risks evaluating tactical decision-making speed rather than the strategic adaptability required under Goldwater-Nichols framework.
Share your assessment in the comments below.