Where Judgment Lives: Thresholds, Targets, and the Upstream Turn in AI-Assisted Warfare

Erin Zhang

Volume 2 • Issue 2

Introduction

Every day, artificial intelligence systems make classification decisions: which emails
belong to spam, which transactions pose fraud risks, which patients are marked as high-risk, and which content will appear in the information flow of social media. Although these decisions may seem unrelated to each other, they all rely on the same underlying mechanism – the classification threshold, which is the boundary that distinguishes one category from another. In the vast majority of cases, this threshold is considered a common technical parameter that reflects a statistical trade-off between false positives and false negatives, rather than an ethical choice. However, the same mechanism also exists in military artificial intelligence systems. When the threshold no longer determines what we read or buy, but who will be recognized as a legitimate military target, technology itself has not changed, but its social and ethical significance has undergone fundamental changes. When a military artificial intelligence system determines that a person constitutes a legitimate target for attack, this decision does not occur at the moment when the operator faces the screen and reviews the system results, but rather earlier – it is completed through the selection of model training data, which behavioral features are defined as “evidence”, and most importantly, where the classification threshold between “possibly combatants” and “possibly non combatants” is drawn.

Current discussions on AI-assisted warfare largely position meaningful human judgment as the final review step: when operators face AI-generated target suggestions, they decide whether to approve the action at this point. Thus, human control is also considered to
exist within this final decision-making process – or precisely where it fails. This paper argues that such an understanding actually misplaces the true decision-making point. The decisive human judgment does not occur after deployment but is already completed during the model design phase. The selection of training data, feature engineering, and particularly the setting of classification thresholds collectively determine who can first enter the system’s field of view and be identified as a legitimate military target. In other words, when human reviewers finally see a name, a risk score, or a target recommendation, those truly decisive judgments have already been completed and frozen in the technical architecture of the system. Therefore, the core research question that this article aims to answer is: at which stage in the AI assisted warfare process does meaningful human judgment occur? How does the setting of classification threshold transform ordinary statistical classification into actionable military decisions with real military consequences?

To answer this question, analysis needs to be conducted at two levels simultaneously. Firstly, it is the technical aspect of classification systems. Threshold is a common and extremely common design in almost all artificial intelligence classification models, widely used in daily AI systems such as spam filtering, credit scoring, fraud detection, and medical diagnosis. Secondly, there is the special social context of war. When the same classification mechanism is applied to military target recognition, the technical logic itself remains almost unchanged, but its social and ethical significance undergoes fundamental changes. The classification boundaries that originally only determined recommendations, screening, or risk levels now begin to determine who can become the “target” in the system, and further affect military operations and even life and death outcomes in reality. Based on the above thinking, this article will argue on four levels. Firstly, this article will explain that the classification threshold is a universal and everyday technical mechanism in all artificial intelligence classification systems, and is not unique to military artificial intelligence. Secondly, this article will analyze how the social and ethical significance carried by the threshold changes when the same classification mechanism enters a war scenario, and why war transforms an originally technical parameter into a political decision. Subsequently, this article will combine Lavender and Project Maven cases to demonstrate that in human-machine collaborative target recognition systems, human review has gradually evolved into a procedural review rather than a truly independent judgment. The truly decisive value judgment has actually been shifted upstream to the model design stage. Finally, this article proposes that if human judgment has moved forward with model design, then the focus of AI governance and regulation in warfare should also be adjusted accordingly. Compared to simply emphasizing the “human in the loop” after deployment, future governance should focus more on the model development phase, including the construction of training data, feature engineering, and the setting of classification thresholds, because it is in these design decisions that statistical probability is transformed into institutionalized actionable military decisions.

Literature Review

“

The threshold is not a neutral adjustment knob. It indicates a decision that someone once made: how much uncertainty a war is willing to transform into certainty about who will die.

The current research on artificial intelligence and warfare mostly revolves around a common problem: how humans supervise when artificial intelligence systems generate suggestions. Among them, Paul Scharre proposed one of the most influential theoretical frameworks currently available. He believes that in military decisions involving life and death, it is necessary to always maintain “meaningful human control” (Scharre, 75) and be wary of weapon systems that can autonomously identify and attack targets without human involvement in the final decision-making process. Scharre’s argument provides an important foundation for understanding human control in military artificial intelligence, but its analysis also implies a premise: truly ethical judgments occur at the point of firing or the final approval of action. In contrast, the design stages such as the model itself, training process, and parameter calibration are considered as the completed technical background, rather than the field where the judgment itself occurred. This article will further advance this discussion based on Scharre’s work, arguing that true human control began earlier than he had imagined. If an artificial intelligence system compresses the real world into a list of “probable targets” through a model before a human operator looks at the screen, then the real control over this list does not come from the final reviewer, but from the designer during the model construction phase.

The concept of “moral crumple zone” proposed by Madeleine Clare Elish provides another important theoretical perspective on this issue. Elish pointed out that in highly automated systems, even if human operators actually have extremely limited control capabilities, responsibility often falls on them, just as the crumple zone of a car absorbs impact through its own deformation to protect passengers inside the car (Elish, 3). Elish’s focus is on the redistribution of responsibility. This article further extends this idea to the reassignment of judgments. In other words, not only is the responsibility constantly focused on the ‘human in the loop’, but true control is gradually moving away from them. The deeper decision-making process itself has already migrated upstream to before model deployment, determined by design choices such as training data, feature engineering, and model parameters.

Bowker and Star provide a deeper theoretical foundation for STS in this article. They pointed out that classification systems have never been neutral. Once established, they gradually evolve into an invisible infrastructure, to the point where the political and ethical labor involved in constructing these classification systems gradually disappears from people’s view (Bowker and Star, 325). Louise Amoure further extends this viewpoint to the field of algorithmic governance. She believes that machine learning systems can only generate partial descriptions about individuals, and ethical responsibility does not only exist in the results output by the algorithm, but also in the inevitable conditions of locality and opacity during the algorithm establishment process. Lucy Suchman further pointed out that the agency between humans and machines is not an objective existence, but a constantly constructed arrangement in concrete practice. In other words, the boundary between what is defined as “human decision” and what is defined as “machine decision” is itself a design choice (Suchman, 15). Similar logic is also reflected in Esqueda’s research on AI Companions. Esqueda pointed out that reinforcement learning systems designed to maximize user satisfaction are not just about responding to the user’s reality, but rather creating a conforming “symptomatic echo chamber” in the process of continuously reinforcing the user’s existing viewpoints, ultimately leading the user to mistakenly believe that this feedback truly reflects themselves (Esqueda). In other words, a seemingly “reflective reality” technological system is actually constantly reshaping reality itself through its design choices. Based on the above researches, it can be found that they all point to an important conclusion, which is also the core viewpoint that this article hopes to further develop: the Lavender and Project Maven discussed in this article are not the real research objects of the paper, but only empirical cases used to illustrate this theoretical proposition. The real focus of this article is not on whether a military AI system has problems, but on how model design can use techniques such as classification thresholds to transform continuous statistical probabilities into action decisions with real military consequences.

Technical Background

Every classification system, however sophisticated, performs the same basic operation: it takes an input, converts it into a set of measurable features, and produces a score expressing the model’s confidence that the input belongs to some category of interest. Taking Support Vector Machine (SVM) as an example, it separates data points of different categories as much as possible by finding an optimal decision boundary. For new input samples, the system predicts their category based on their location on the decision boundary and distance from the boundary, and generates corresponding confidence scores. However, confidence scores themselves do not automatically generate decisions. The system still needs to set a classification threshold: when the score exceeds this threshold, the system outputs ‘Yes’; when it is below this threshold, output ‘No’ (see fig. 1).

Fig. 1. Classification threshold converts continuous probability scores into binary decisions. Different threshold settings produce different trade-offs between false positives and false negatives. Adapted from Google Machine Learning Crash Course, Google Developers.

This threshold is not ‘discovered’ by the algorithm itself, but is set by the model developer or deployer according to actual needs. The setting of a threshold essentially means balancing between two types of errors. On the one hand, it is a false positive, where the system
incorrectly identifies an object that does not belong to the target category as the target. On the other hand, there is a false negative, where the system fails to recognize objects that truly belong to the target category. Raising the threshold can reduce false positives, but it will increase false negatives; lowering the threshold is exactly the opposite. There is no threshold that can simultaneously minimize two types of errors. This trade-off actually exists widely in various daily artificial intelligence applications. For example, in systems such as spam filtering, credit scoring, medical screening, and financial fraud detection, setting thresholds is essentially an institutional judgment: which type of error is an organization willing to bear the cost of. Taking disease screening as an example, if the threshold is set lower, the system will mark more patients as “possibly ill”, thereby increasing the detection rate of real cases. However, it will also bring more false positives, requiring a significant amount of additional resources for follow-up examinations. On the contrary, if the threshold is raised, it can reduce false positives but may miss patients who truly need treatment. Therefore, no clinical doctor would simply understand this threshold as a technical detail. On the contrary, it is often seen as a policy level decision and its reasonable boundaries are continuously discussed by relevant medical professional organizations. For classification systems in other fields, the same level of scrutiny should be given, even if these thresholds are only hidden in the model card or configuration file and hardly noticed by anyone except the engineering team.

This seemingly ordinary technical fact forms the basis of all the arguments presented in this article. The classification threshold is not only the location where statistical uncertainty is transformed into binary classification results, but also the key node where probabilistic knowledge is transformed into institutionalized actionable certainty. Once the threshold is set, uncertainty is no longer just a statistical calculation result, but is institutionalized through organizational processes, military norms, and specific actions. Therefore, regardless of whether this decision is explicitly recognized or not, it always implies that someone has decided where this transformation should occur, and this decision itself is a political and ethical judgment.

Same Algorithm, Different Context

RETURN TO THE ISSUE

The Synthetic Battlefield

Volume 2 • Issue 2

The classification mechanism discussed earlier does not fundamentally alter its technical logic when transitioning from the commercial domain to the military domain. A model designed to evaluate loan applicants and one designed to assess suspected combatants can share the same underlying architecture, the same concept of decision boundaries, and the same threshold logic for classification. What truly changes is not the algorithm itself, but the cost incurred by false positives. When the threshold of the spam filter is set too loosely, a normal email may be mistakenly classified as spam – although this can cause inconvenience, it can still be restored. When the threshold calibration of the credit scoring model is not correct, a qualified applicant may lose the loan opportunity – this is a real harm, but it still occurs in an institutional system with appeal, supervision, and subsequent correction mechanisms. However, when the threshold of the target recognition model is set too loosely, a false positive may mean that a non combatant person has been mistakenly added to the kill list. In these three scenarios, the statistical calculation process is completely consistent, what truly changes is the ethical weight carried by the same mistake.

It is precisely in this rupture where the technological logic remains unchanged while the social consequences undergo significant changes that Bowker and Star’s argument that “classifications and standards give advantage or they give suffering” (Bowker and Star, 6) is no longer just a metaphor, but has become a concrete reality. Similarly, Amore’s viewpoint that algorithms can only provide partial descriptions is no longer just a reminder at the epistemological level, but has become a practical question about who gets to keep living. The issue is further exacerbated by military and intelligence agencies, as they often view model outputs as intelligence about the real world rather than its true essence – a statistical guess. This speculation is based on the labels formed during the training process and is deeply influenced by the classification threshold setting, expressing only the probability of a certain behavior pattern being associated with a predetermined label. The number attached to a name may seem precise, but in reality it encodes an earlier policy decision that occurred elsewhere: how much statistical uncertainty an organization is willing to convert into institutional certainty.

The recruitment algorithm can further highlight this difference. When the threshold of the recruitment system is set too loosely, a job seeker who does not meet the requirements may enter the interview process – an error that the recruitment committee can tolerate and correct in subsequent stages. However, when the target recognition system operates with the same relaxed threshold, the corresponding errors will not be corrected in the next stage, but will be directly executed. These two types of systems can share the same code, the same logistic function, and even be developed by the same group of engineers switching between commercial and defense projects. But what they cannot share is the reversibility of errors generated by thresholds. Therefore, once the classification threshold is deployed in military scenarios, it can no longer be understood as a purely technical trade-off that exists within engineering. The same technological choice may only be a negligible error in the recruitment process, but in the context of war, it becomes a critical moment when statistical speculation ultimately solidifies into institutional fact.

Lavender and Project Maven: Empirical Evidence

Yuval Abraham’s report on the Israeli military’s Lavender system, based on interviews with six intelligence officials who had worked on the system, provides a vivid example of the aforementioned mechanism. According to the report, Lavender used hundreds of features extracted from surveillance data to rate approximately 2.3 million residents of Gaza on a scale of 1 to 100, indicating the likelihood of their association with Hamas or Palestinian Islamic Jihad organizations. These features include WhatsApp group membership, behavior patterns of changing phones or addresses, and other behavioral signals used as proxies for armed groups. In the early stages of the war, the system generated a potential target list of approximately 37,000 people. Abraham’s report shows that the threshold used to determine whether a person will be listed will be adjusted up or down based on action pressure in order to generate more or fewer target names. According to reports, human intelligence officers responsible for reviewing lists typically only spend about 20 seconds reviewing each file for lower priority targets. In many cases, they only approve further action after confirming that the object marked by the system is male. Subsequently, a supporting system called “Where’s Daddy” will continue to track individuals who have been marked until they enter the family residence, at which point they may be authorized to launch a strike. No matter how people evaluate this specific project, the structural issues that this article focuses on are narrower and more enduring than any specific fact of a war: a twenty second review is not the place where meaningful judgments about “who should be considered a target” truly occur (Abraham). The true judgment is already completed at an earlier stage – when someone decides which features can be considered evidence, how many confidence scores are sufficient to cross the boundary of the ‘target’, and how far this boundary can be moved under action pressure, the judgment has already occurred.

Fig. 2. Tactical airborne drone representative of the surveillance platforms whose imagery is processed by AI-enabled computer vision systems such as Project Maven. Source: U.S. Department of Defense.

Project Maven provides an inspiring comparative case precisely because it was originally designed around an intention opposite to Lavender. In 2017, Bob Work, then Deputy Secretary of Defense of the United States, launched the project under the name of the Algorithmic Warfare Cross Functional Team (Work). Maven applies machine learning object detection technology to unmanned aerial vehicles and satellite imagery, and is clearly positioned as a decision support tool rather than an autonomous object recognition system. Before any combat decision is made, human analysts are still responsible for reviewing objects marked by artificial intelligence. However, the subsequent development of the project has shown that even if a system is explicitly designed with the goal of retaining downstream human review, it still cannot fully respond to the issues arising from upstream classification selection. Maven was later transferred to the National Geospatial Intelligence Agency of the United States and expanded into a formal project used by multiple combat commands. At the same time, Google also withdrew from related contracts in 2018 after employees opposed the use of its computer vision research in the military field. These developments raise several questions that cannot be solved solely by final human review: what can be defined as an “object of interest”, which training images determine the meaning of this category, and what confidence threshold is sufficient to trigger system labeling and attract the attention of human analysts. Matt Mande and Gregory Allen’s analysis of the project’s evolution further reinforces this viewpoint: Maven has evolved from an image classification pilot project to a platform adopted by almost all combat commands in the United States, constantly absorbing new data types, and has been managed by more than one main contractor since Google withdrew in 2018 (Mande and Allen). In fact, each such transformation is equivalent to a re-calculation of the upstream design problem that this article focuses on: new training data, new features that are likely to appear, and threshold adjustments made to adapt to the new data processing flow. However, the public narrative about the project still focuses on the human analysts who output the review system, rather than the design choices upstream of this review process.

Lavender and Project Maven jointly illustrate a recurring structural pattern: judgments are focused on the design phase, while reviews are compressed into confirmation procedures. Even if these systems are established with different intentions and are subject to different levels of public and legal scrutiny, this model will still recur. As emphasized in the research blueprint of this article, the role of these two cases in this section is only to provide empirical evidence for a general proposition about model design. They are not the core argument of this article, nor should they be understood as: what truly matters is the specific facts in each case, rather than the structural patterns presented by both.

Human Judgement Moves Upstream

The existing research on artificial intelligence and war mainly focuses on how to retain or restore human supervision after model generation of suggestions. Scharre’s “meaningful human control” and Elish’s “moral crush zones” both consider the moment when human review occurs as the core domain where responsibility and control are in dispute. The original contribution of this article lies in pointing out that this theoretical framework incorrectly locates the location where judgments occur. The ‘human in the loop’ in AI assisted warfare has not disappeared, it only shifts upstream and is embedded in the three specific design stages before deployment, and the current governance discussions have hardly examined these stages.

The first step is the selection of training data. Each model learns what a “target” is from a labeled dataset, and determines which individuals’ past behavior can be used as examples of “combatants”, which is itself an explanatory and controversial judgment, rather than a neutral data collection behavior (Bowker and Star, 24). The second step is feature engineering. Viewing WhatsApp group membership or frequent phone changes as evidence of association with armed groups is a defining behavior disguised as measurement language. Viewing a certain behavioral pattern as a meaningful “feature” is itself a judgment of what the association relationship should look like. Engineers or analysts who make such judgments may never personally review any specific personal files. The third and most significant consequence is the threshold selection placed at the core of this article. Threshold is not just an engineering parameter adjusted to improve accuracy. They are frozen human judgments that transform statistical uncertainty into institutional certainty. When someone sets a score at which a person will be “marked” by the system, they are actually making a decision in advance and on a large scale: to avoid a certain number of false negatives, how many false positives the institution is willing to accept – that is, how many people who should not have been the target are incorrectly marked. This decision will not be made again when an officer spends twenty seconds reviewing the file. It has already been completed in one go upstream and then uniformly applied to every subsequent case downstream.

This redefinition changes how we can honestly understand “human-in-the-loop”. The officer reviewing the Lavender file was not exercising the kind of judgment Scharre wanted to retain. This kind of judgment has already occurred when others choose a scoring system of 1 to 100 points, decide which features can enter the model, and set how high a score a name needs to reach in order to appear on the list. Combining Amore’s attention to the opacity and locality of algorithm descriptions, as well as Suchman’s argument that the decision boundary between humans and machines is not discovered but constructed in practice, the upstream turn proposed in this article means that the distinction between “determined by humans” and “determined by machines” was created in advance by those who wrote this boundary into the system design. If you are still searching for judgment in front of the downstream review screen, you are actually looking for something that has already been transferred to another location.

This viewpoint also reveals a limitation of Elish’s crumple-zone metaphor, although this metaphor is still very useful. The crumple zone of a car is a physical structure that is fixed during the manufacturing stage, and it absorbs impact in a predictable manner regardless of where the collision occurs. However, the moral crumple zone described in this article is not so passive. The problem is not just that the responsibility will fall on the human operator afterwards, but that the operator has not been placed in a position to exercise the judgment required of them from the beginning, because this judgment has already been consumed in the selection of data, features, and thresholds, and has been completed long before the operator starts duty. This distinction is crucial for the attribution of responsibility. If it is judged that the decision has shifted upstream, then requiring the auditor to take responsibility for a wrong target identification decision wrongly identifies the person who actually made the key decision. This is like blaming cashiers for mis-pricing products in a store: it mistook the person who executed the predetermined result for the person who set the price.

Conclusion

If meaningful human judgment in artificial intelligence assisted warfare mainly exists in training data selection, feature engineering, and threshold setting, then the governance framework that only regulates deployment and review stages – such as reviewing whether humans have “look at” system recommendations or applying the principle of “meaningful human control” only during weapon launches – is actually reviewing the wrong moments.

This does not mean that human reviews during the deployment phase are worthless, but rather that people overestimate the substantive significance of such reviews compared to design decisions that occurred earlier and elsewhere. Effective governance must move upstream to correspond to the location of the true occurrence. This means that the system must disclose and explain the trade-off between the false positives and false negatives it has chosen. The source, formation process, and labeling method of training data should be considered as policy issues that can be reviewed, rather than proprietary technical details. The selection in feature engineering – which behavioral signals can be allowed to serve as alternative indicators for lethal categories – should also be subject to the same level of scrutiny currently reserved only for final firing decisions. Specifically, this may mean requiring the military to disclose its threshold and the corresponding error rate before the system is deployed in combat. Even if this information needs to be submitted in a confidential form, it should be able to be reviewed by supervisory agencies. This is similar to clinical trials requiring preregistration of decision criteria for drugs before they actually come into contact with patients.

The above proposition does not mean that upstream governance is easy to implement, nor does it necessarily mean that it can prevent any specific decision recorded in the Lavender and Project Maven cases mentioned earlier. This article proposes a more limited proposition: where should the analytical gaze be directed. As Esqueda described in discussing artificial intelligence companions, the feedback loop of conformity is not just about reproducing the world, it also shapes what can become visible in the world, and the threshold is one of the most significant and least scrutinized locations when this shaping occurs.

The threshold is not a neutral adjustment knob. It indicates a decision that someone once made: how much uncertainty a war is willing to transform into certainty about who will die. This decision only needs to be made once, but it will be widely applied to subsequent cases. This article argues that the human judgment that is constantly searching for errors in current discussions actually exists here. Identifying the location where this decision occurred is the first step that must be taken to govern it.

Works Cited

Abraham, Yuval. “‘Lavender’: The AI Machine Directing Israel’s Bombing Spree in Gaza.” +972 Magazine, 3 Apr. 2024,
https://www.972mag.com/lavender-ai-israeli-army-gaza/. Accessed 23 July 2026.
 
Amoore, Louise. Cloud Ethics: Algorithms and the Attributes of Ourselves and Others. Duke University Press, 2020.

Bowker, Geoffrey C, and Susan Leigh Star. Sorting Things Out: Classification and Its Consequences. Mit Press, Cambridge, 1999.

Elish, Madeleine Clare. “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction.” Engaging Science, Technology, and Society, vol. 5, Mar. 2019, pp. 40–60, https://doi.org/10.17351/ests2019.260.

Esqueda, Alexa. “AI Companions: The Echo Chamber of Sycophancy – Simulation & Society.” Simulation-And-Society.Org, 2025,
https://simulation-and-society.org/ai-companions-the-echo-chamber-of-sycophancy/. Accessed 22 July 2026.

Google Developers. Machine Learning Crash Course: Thresholds and the Confusion Matrix. Google,
https://developers.google.com/machine-learning/crash-course/classification/thresholding. Accessed 23 July 2026.

Lucille Alice Suchman. Human-Machine Reconfigurations : Plans and Situated Actions. Cambridge Univ. Pr, 2007.

Mande, Matt, and Gregory C Allen. “What Is Maven Smart System, and What Does It Do?” Csis.Org, 2026,
https://www.csis.org/analysis/what-maven-smart-system-and-what-does-it-do. Accessed 23 July 2026.

Scharre, Paul. Army of None: Autonomous Weapons and the Future of War. W.W. Norton & Company, 2019.

“Tactical Drone.” Photograph by John Smith. U.S. Department of Defense, 24 Aug. 2020, VIRIN 200824-D-ZZ999-002C.JPG,
https://www.defense.gov/Multimedia/Photos/igphoto/2002485623/. Accessed 23 July 2026.

Work, Bob. “Department of Defense, Memo: Establishment of an Algorithmic Warfare Cross-Functional Team (Project Maven), April 26, 2017.” NSArchive, 26 Apr. 2017,
https://nsarchive.gwu.edu/document/18583-national-security-archive-department-defense. Accessed 22 July 2026.