A cross-domain framework for modelling and evaluating detection systems against threats designed to evade them, with hypersonic tracking as the canonical case

Abstract

This paper treats four apparently unrelated detection problems — hypersonic objects, small uncrewed aircraft and swarms, adaptive and AI-mediated digital threats, and biological signals — as a single class, unified not by their physics but by a shared adversarial structure: in each, the thing to be detected is weak, embedded in variable clutter, present in growing numbers, or actively adapting, and in each the decision-relevant question is not whether a sensor can produce a detection in a favourable demonstration but whether a system can sustain timely, trustworthy target custody under stress. The paper’s contribution is a cross-domain evaluation framework — a five-layer directed sensing-to-decision chain, assessed across a scenario matrix rather than by a single score — and a taxonomy that maps four threat families onto that framework. Hypersonic detection is developed as the canonical worked case. The paper is methodology-forward and open-source throughout; it makes no claim about classified-system capability, presents no operational detail on building, optimising, or evading any threat, and treats the biological branch strictly at the level of sensing and surveillance. Its central inference, committed to in print, is that the binding constraint on detection performance is increasingly the absence of end-to-end, condition-stratified evaluation, and that a nominal receiver-operating-characteristic curve is not an adversarial one.

1. Introduction: a class of problems, not a collection of them

The detection of a hypersonic glide vehicle, the detection of a small drone in a swarm, the detection of an adversarially perturbed input to a machine-learning classifier, and the detection of a biological attack through a public-health surveillance network are ordinarily treated as four separate engineering problems, addressed by four separate communities using four separate vocabularies. This paper argues that they are better understood as four instances of a single class, and that seeing them as a class yields an evaluation discipline that none of the four communities reliably applies on its own.

The unifying feature is in the paper’s title. In each of these problems, the thing to be detected does not want to be seen — either because an adversary has designed it to be hard to detect, or because its very nature makes it faint, fleeting, numerous, or mutable. A hypersonic object presents a weak and variable signature against a cluttered background and transits its engagement window in a very short time. A small uncrewed aircraft is individually faint and, in a swarm, one of many. An adaptive digital threat changes its observable characteristics specifically to defeat the detector that hunts it. A biological attack may be indistinguishable, in its early signal, from the ordinary noise of public health. What these share is not a sensor, a physics, or a domain. It is a structure: detection under conditions of weak signal, variable clutter, growing target population, or deliberate adaptation.

The consequence of this shared structure is a shared failure mode in how detection is evaluated, and it is this failure mode the paper exists to correct. Across all four domains, the tempting evaluation is the favourable demonstration: the successful flight test, the high benchmark accuracy, the assay that detects the analyte, the sensor that produces a detection under good conditions. And across all four, that evaluation is systematically misleading, because operational utility does not depend on whether a detection can be produced under favourable conditions. It depends on whether timely, correctly associated, trustworthy target custody can be sustained under unfavourable ones. The gap between those two things — between the demonstration and the operational reality — is where detection systems fail, and it is precisely the gap that a favourable demonstration is structured not to reveal.

This paper proposes a framework for closing that gap. It is a companion, in method and spirit, to the treatment of stress-extended verification and validation for commercial-off-the-shelf components elsewhere in this volume; both papers hold that performance is a property of a system operating under conditions rather than of a component observed in isolation, and both insist that honest evaluation must be organised around the conditions under which a system is expected to fail rather than the conditions under which it happens to succeed. Where the companion paper addresses components under environmental and operational stress, this paper addresses detection systems under adversarial and clutter stress, and extends the discipline across four threat domains.

A word on what this paper is and is not. It is a methodological framework, developed from open-source technical, governmental-analytical, and public-health literature. It is not a systematic review, and it does not claim comprehensive coverage of any programme. It makes no claim about the capability of any classified system. It contains no operational detail that would assist in building, optimising, acquiring, or evading any of the threats it discusses; consistent with the evidence scan on which it draws, it excludes vehicle design, countermeasure tactics, pathogen engineering, and acquisition, and it treats the biological domain strictly at the level of sensing and surveillance. Its subject is evaluation — how to know how well a detection system actually works — and evaluation is a defensive discipline.

2. The evaluation backbone: five layers that are usually collapsed

The core of the framework is a separation. Detection performance is routinely reported as though it were a single quantity — a detection probability, an accuracy figure, a hit rate — when it is in fact the composite output of at least five distinct layers, each with its own failure modes, each capable of degrading independently, and each of which a favourable demonstration tends to collapse into the others. A defensible evaluation separates them and reports each. The five layers are sensing, detection-and-classification, tracking-and-fusion, decision-timeliness, and robustness.[1][2]

The sensing layer asks what physical or digital signature is actually observable, at what range or signal quality, and against which background and clutter conditions. This is prior to any question of algorithm or threshold; it concerns whether the information required for detection is present in the observable signal at all, and how that presence varies with geometry, weather, occlusion, and the state of the target. A sensing layer that is adequate under favourable geometry may carry almost no exploitable information under unfavourable geometry, and an evaluation that does not vary the geometry will not discover this.

The detection-and-classification layer asks at what threshold the system declares an alert, and — crucially — how missed detections, false alerts, and misclassifications change as that threshold moves. This is the layer at which the receiver-operating-characteristic curve lives, and it is the layer most often mistaken for the whole. A threshold sweep here is necessary. It is also, as the paper argues below, radically insufficient, because it captures the local detection-versus-false-alarm trade-off while saying nothing about time, track quality, or the downstream cost of a false response.

The tracking-and-fusion layer asks how rapidly an alert is associated into a confirmed track, and how often that custody breaks, swaps between targets, duplicates, or persists after the target has gone. This layer is where a detection becomes a usable track, and its failure modes — fragmentation, identity switching, false tracks, ghost persistence — are invisible to any evaluation that stops at the single-detection level. Fusion belongs here: combining repeated or heterogeneous measurements can raise effective detection performance, but only if the evaluation simultaneously represents the false tracks, association errors, latency, and cross-sensor dependence that fusion introduces. Fusion is not free, and an honest fusion evaluation prices it.[3][4]

The decision-timeliness layer asks whether the alert or track arrives early enough for the specific downstream decision it is meant to inform, after accounting for communication and fusion latency. A detection that is correct but late is, for many purposes, a detection that has failed, and lateness is a property not of the sensor alone but of the whole chain from observation through fusion to the point of decision. This layer forces the evaluation to specify the decision the detection serves and the time budget that decision imposes, and to measure the system against that budget rather than against an abstract notion of accuracy.

The robustness layer asks how every one of the preceding measures changes when conditions are varied — when background, weather, target density, sensor availability, data distribution, or the adversary’s behaviour departs from the nominal. Robustness is not a separate metric but a transformation applied to all the others: it is the difference between a system’s performance under the conditions of its demonstration and its performance under the conditions of its use. It is, in the argument of this paper, the layer that matters most and is measured least.

3. The constrained trade-space and the scenario matrix

The five layers share a common currency, which the framework calls the constrained trade-space. Across every domain, the decision-relevant quantities are the same small set, and the discipline the framework imposes is that they be reported together and under defined operating conditions, never singly and never unconditioned. They are: detection probability (equivalently sensitivity or recall); false-alert burden (equivalently false-positive rate, or its complement in precision); time to alert or track confirmation; localisation and classification error; and continuity of the track. A number drawn from any one of these in isolation — a high detection probability, a low false-alarm rate, a fast alert — is not an evaluation. It is a fragment of one, and often a misleading fragment, because the quantities trade against one another and the trade is the point.[5][6]

The reason they must be reported together is that each can be improved at the expense of the others, and a system tuned to excel on one metric under demonstration conditions may be quietly catastrophic on another under operational conditions. A detector can achieve a high detection probability by lowering its threshold until its false-alert burden overwhelms the downstream decision-maker. It can achieve a low false-alarm rate by raising its threshold until it misses the targets that matter. It can achieve a fast alert at the cost of track quality, or high track quality at the cost of timeliness. Only the joint report, under stated conditions, reveals whether the system is actually fit for its decision.

This is why the receiver-operating-characteristic curve, indispensable as it is, cannot be the whole of an evaluation. The ROC curve expresses the local trade between detection and false alarm at the classification layer. It is silent on time, on track quality, and on the cost of a false response — the very quantities on which operational utility turns. A public-health analogue makes the same point in a different vocabulary: the framework for evaluating public-health surveillance holds that evaluation should compare timeliness against the balance of sensitivity and predictive value, define the desired sensitivity, specificity, and event scale in advance, and measure detected, missed, late-detected, and false alarms as a routine part of the workflow rather than reporting a single headline figure.[7] The domains differ; the discipline is identical.

From this follows the framework’s central experimental prescription: the output of a detection evaluation should be a scenario matrix, not a single aggregate score. Results must be stratified — at minimum by signal-to-clutter condition, target density, geometry and occlusion, sensor availability, and time-to-decision budget, and, where an adversary is present, by adversary capability. Within each stratum the evaluation should report uncertainty, as confidence intervals or bands, rather than a point estimate. And the evaluation should refuse two tempting summaries: it should not offer a high average across the matrix as evidence of adequacy, because the average conceals the strata where the system fails, and it should not offer a single successful test or flight event as evidence of end-to-end performance across the matrix, because one favourable point is not a distribution. The public-health guidance again anticipates the method, recommending repeated scenario simulations to build operating-characteristic curves across a range of conditions while candidly noting the limits of simulation in representing real events.[8]

Epistemic layer — confirmed from open sources (1 of 3): that detection performance is a property of a system operating under conditions rather than of an isolated sensor, and that fusion improves performance only subject to false-track, latency, and correlation penalties, is established across the cited multi-sensor fusion and cluttered-environment tracking literature.[9][10]Epistemic layer — confirmed (2 of 3): that ROC-style threshold sweeps are necessary but insufficient, and that timeliness must be weighed against sensitivity and predictive value, is established and is expressed identically in the public-health surveillance-evaluation framework.[11] The five-layer decomposition and the scenario-matrix prescription, as a single cross-domain framework, are the author’s contribution (1 of 2), built on these confirmed foundations.

4. Canonical worked case: hypersonic detection, tracking, and handoff

The framework is best seen at work, and hypersonic detection is the canonical case because it exhibits the whole class in concentrated form: a signal that is weak and variable, a background that is cluttered, a decision window that is short, and — in the adversarial framing that defines the class — a target whose physics and trajectory make sustained custody genuinely hard. The open technical literature is clear that the hard problem is not initial discovery but continuous detection, discrimination, and handoff: the challenge is to acquire the object, hold custody of it, discriminate it from clutter and decoys, and pass a track of sufficient quality to the point of decision, continuously, across the whole engagement.[12] This paper treats the hypersonic case strictly at that architectural and evaluative level; it contains nothing about the design, performance, or defeat of any vehicle, and draws only on the open, published detection literature.

The physical and geometrical constraints that make the problem hard are, at the level the open literature describes them, a matter of sensing physics: atmospheric absorption and scattering attenuate signals, line-of-sight is obstructed by the curvature of the Earth and by terrain, and the physics of high-speed flight complicate the observable signature. It is precisely because no single sensor or vantage overcomes all of these simultaneously that the reviewed literature motivates adaptive, multi-platform, and multimodal sensing — not as a preference but as a structural necessity imposed by the constraints.[13] A published small-satellite concept from 2023 illustrates the measurement-level choices such an approach entails: a cryogenically cooled mid-wave infrared imager for the primary signature, complementary visible imagery for cross-validation, spectral filtering informed by atmospheric models, and onboard processing, with optical and radiometric validation planned. It is valuable precisely as an engineering design example — and, as its own authors are careful to note, it is a design example rather than evidence that a full architecture has achieved operational custody.[14] The distinction between a demonstrated component and a validated architecture is exactly the distinction this paper’s framework exists to preserve.

The publicly described architecture for space-based hypersonic tracking makes the fusion logic concrete, and it maps directly onto the framework. In the architecture as described in open government-analytical sources, wide-field-of-view sensors provide persistent initial detection and cueing, while more sensitive medium-field-of-view sensors refine the resulting track; the stated objective is infrared detection against Earth-background clutter and the provision of track data of a quality sufficient for downstream engagement decisions.[15] This description directs the evaluation toward four linked transitions, and the framework’s discipline is to treat each as a distinct object of measurement rather than to credit the whole to a single successful alert. The four transitions are: wide-field detection of the object; delivery of the cue from the wide-field to the medium-field layer; narrow-field acquisition and track refinement; and delivery of the refined track to command and control.

Each of these four transitions requires, in the framework, its own probability of success, its own latency distribution, and its own track-error or covariance measure — and the end-to-end metric for the architecture must condition on all four jointly rather than reporting any one in isolation. A system that detects reliably in the wide field but loses objects at the cue-delivery seam has not achieved tracking; a system that acquires well in the narrow field but delivers its track to command and control too late for the decision it serves has not achieved tracking either. The single most important discipline the framework imposes on the hypersonic case is that a successful first alert is not completed tracking, and that an architecture must be scored on the full chain from detection through delivery, at each seam.

The joint end-to-end metric for the four-transition hypersonic architecture can be stated in the framework’s terms without any recourse to system-specific values. Let each of the four transitions — wide-field detection, cue delivery, narrow-field acquisition and refinement, and delivery to command and control — carry, within a given scenario, a probability of success, a latency distribution, and a track-error measure. The architecture achieves useful custody, for a given downstream engagement concept, only when all four transitions succeed, when their summed latency falls within the decision window that concept imposes, and when the resulting track error meets that concept’s accuracy requirement. The end-to-end metric therefore conditions jointly on success at every seam, on the convolution of the four latency distributions lying within budget, and on the propagated track error meeting threshold — and it must be reported separately for each engagement concept, since the open assessment itself distinguishes the tracking accuracy needed for different remote-engagement concepts. A single architecture may satisfy the joint metric for a less demanding concept and fail it for a more demanding one, and an evaluation that reports one unconditioned number conceals exactly that distinction.

This formal statement makes visible why the seams, and not the sensors, are the right primary objects of evaluation for the hypersonic case. The individual sensing performances — how well the wide-field layer detects, how well the medium-field layer refines — are necessary inputs, but the architecture’s real performance lives in the conditional probabilities at the transitions between them and in the convolution of their latencies. A programme that measures and reports its sensor performances, however rigorously, while leaving the seam probabilities and latency distributions unmeasured, has characterised the parts of its architecture that were never the hard problem and left uncharacterised the parts that were. The framework’s contribution to the hypersonic case is to move the primary measurement from the sensors to the seams — and the joint metric is the instrument that forces that move.

This points to the deepest observation the framework makes about the hypersonic case, and it generalises to the whole class. The central uncertainty in such an architecture is not whether it is conceptually coherent — the wide-field-cue-then-refine logic plainly is — but how its performance degrades at the seams between its layers. The same open government assessment that describes the architecture also notes the dependence of the sensitive medium-field layer on cueing from the wider layer, and distinguishes the tracking accuracy required for different downstream engagement concepts.[16] The framework’s prescription follows directly: a model of such a system must treat sensor handoff, communications links, and track correlation not as invisible, always-available infrastructure but as stochastic nodes with their own failure probabilities and latency distributions. The seams are where the architecture is weakest and where evaluation is thinnest, and the framework’s contribution to the hypersonic case is to make the seams the primary objects of measurement.

Epistemic layer — confirmed from open sources (3 of 3): that the open hypersonic-detection problem is one of continuous detection, discrimination, and handoff rather than initial discovery, and that the publicly described wide-field-cue / medium-field-refine architecture depends on cueing across a seam, is established in the cited technical and government-analytical literature.[17][18][19] The treatment of the four transitions as jointly conditioned evaluation objects, and of the seams as stochastic failure nodes to be modelled explicitly, is the author’s framework applied to the case (2 of 2). No classified capability is asserted, and no vehicle-level detail is present.

5. Taxonomy branch I: small uncrewed aircraft and swarms

The first branch of the taxonomy varies one axis of the trade-space — target density — and thereby exposes a different set of the framework’s layers to stress. Where the hypersonic case is the few-fast-faint version of the problem, small-uncrewed-aircraft detection is the many-faint-target version: individually weak signatures, present in numbers, in a low and cluttered air picture. The open literature establishes that the available sensing modalities — radar, radio-frequency, electro-optical and infrared, and acoustic — are complementary rather than interchangeable, each strong where others are weak. A signature study examined acoustic, passive and active optical, and frequency-modulated continuous-wave radar sensing and concluded that their differing strengths can be combined into a complementary network; a European Commission Joint Research Centre overview of counter-uncrewed-aircraft techniques likewise identifies radar, RF, EO, and acoustic sensing and argues for real-time multi-sensor fusion precisely because heterogeneous sensors and platforms complicate detection, tracking, and identification.[20][21]

The branch’s distinctive contribution to the framework concerns the fusion-and-tracking layer under density, and it introduces a metric the hypersonic case does not foreground: throughput. Fusion in this domain must be evaluated against both accuracy and processing rate, because the two trade against each other in ways that a single-target evaluation never reveals. A systematic study of infrared-and-visible camera integration for small-uncrewed-aircraft detection found that decision-level fusion improved recall and precision but reduced frame rate, while pixel-level fusion improved neither.[22] The lesson the framework draws is precise and general: in dense-target detection, the evaluation must report detection and classification accuracy alongside frame rate, track-confirmation delay, identity-switch and association errors, and the maximum number of simultaneous tracks the system can hold. A fusion approach that improves accuracy while collapsing throughput may be worse, not better, against a swarm.

The framework therefore adds a specific experimental requirement for the swarm branch: load-scaling curves. The performance of a detection-and-tracking system against one target, or a handful, tells one almost nothing about its performance against a saturating density, because the association and tracking layers degrade non-linearly as the number of simultaneous targets rises — identity switches multiply, tracks fragment, and throughput falls. An evaluation that extrapolates from single-target or few-target trials to swarm conditions is making exactly the error the framework is built to prevent: mistaking a favourable demonstration for operational performance. The correct output is a family of curves showing how each trade-space quantity degrades as target density increases, stratified as always by clutter and geometry. The swarm branch, in short, stresses the tracking-and-fusion and decision-timeliness layers under load, and its evaluation must be built to expose that stress rather than to average it away.

The load-scaling requirement deserves one further specification, because it is where the swarm branch most sharply departs from single-target intuition. The quantities that degrade under density do not degrade uniformly, and the framework asks that their differing degradation rates be measured separately rather than summarised. Track-confirmation delay may rise gently with density while identity-switch rate rises steeply; maximum simultaneous track capacity may hold until a saturation point and then collapse; false-track rate may grow as the association layer, overwhelmed, begins to knit spurious tracks from unassociated detections. Each of these has a different curve, and each curve has a different operational meaning: a system that holds accuracy but saturates its track capacity fails differently from one that holds capacity but loses track identity, and a swarm of a given size may be within one system’s envelope and beyond another’s for reasons a single accuracy figure would never reveal. The branch’s discipline, then, is not merely to produce load-scaling curves but to produce them per quantity, and to identify for each the density at which it ceases to meet the downstream decision’s requirement — the point, in other words, at which the swarm has won.

6. Taxonomy branch II: adaptive and AI-mediated evasion

The second branch varies the most consequential axis of all, and it is the axis from which the paper takes its title most directly: deliberate, intelligent adaptation by the thing to be detected. In the hypersonic and swarm branches the target is hard to detect because of its physics and its numbers; in this branch it is hard to detect because it changes, specifically and responsively, to defeat the detector. This is the domain of adversarial machine learning and adaptive digital threats, and it forces a fundamental revision of what an evaluation must contain.

The revision is this. In a static domain, a conventional held-out test set measures a detector’s discrimination adequately, because the distribution the detector faces in use resembles the distribution on which it was tested. In an adaptive domain that assumption fails, because the adversary is actively moving the distribution to find the detector’s blind spots. The evaluation target must therefore include an explicitly named threat model: what the adversary is able to change, what the detector is able to observe, what the adversary knows about the detector, and the strength and cost of the attempted evasion. Without a threat model, an accuracy figure in an adaptive domain is not merely incomplete; it is measuring the wrong thing, because it characterises the detector against a stationary world that the adversary has no intention of providing.[23]

The open evidence on how far this can go is sobering and is the branch’s empirical anchor. A robustness study that evaluated false-negative-based measures under both static and dynamic adversaries found no examined classifier robust once the adversary sought even a single misclassification.[24] In an Android-malware case study, four nominally strong detection models saw their average accuracy fall from 95.13 per cent to 59.97 per cent under a constructed evasion attack — a collapse from apparent reliability to near-uselessness produced not by any change in the detector but by an adversary adapting to it.[25] The general principle the framework extracts is not that every detector fails in the same way or to the same degree; it is a single sentence that the paper commits to as its sharpest formulation: a nominal receiver-operating-characteristic curve is not an adversarial one. The performance a detector exhibits against a fixed test set is not the performance it will exhibit against an adversary who adapts, and the difference can be the difference between 95 per cent and a coin-flip.

The framework’s prescription for this branch follows. An adaptive-domain evaluation should sweep the adversary’s capability or perturbation budget rather than testing at a single fixed strength; it should include adaptive retesting after any defensive update, because a defence that closes one evasion route may open another; it should measure false negatives and alert burden separately, since an adaptive adversary can attack either; and it should state explicitly which classes of manipulation were not tested, because the untested classes are where the fielded system will be attacked. Crucially, the framework holds that this discipline is not confined to the digital domain. It transfers to physical sensing wherever target behaviour, background manipulation, spoofing, or data-association stress can change the observed signal — which is to say, wherever an adversary can adapt. The adaptive branch is thus not a separate problem from the hypersonic and swarm branches but the general case of which they are constrained instances, and its central lesson — evaluate against an adapting adversary, not a fixed test set — is the lesson the whole class most needs.

Honest limit, stated plainly: the strongest quantitative evidence for adversarial collapse in the cited literature is drawn from digital classifiers — image and malware models — rather than from physical detection systems.[26][27] The paper’s claim that the same discipline transfers to physical sensing is an argued extension, not a demonstrated result, and is offered as such. The transfer is well motivated — spoofing, decoying, and background manipulation are physical analogues of digital evasion — but the paper does not present the digital evidence as if it were physical-sensing evidence, and a reader should treat the physical-domain claim as a reasoned hypothesis to be tested rather than an established finding.

7. Taxonomy branch III: biosurveillance and biological detection

The third branch extends the framework into the biological domain, and it is treated here with a deliberate and stated restriction: it remains entirely at the level of sensing and surveillance. Nothing in this section concerns the characteristics, creation, or acquisition of any biological agent; the subject is exclusively how one detects a biological signal, and how one evaluates a system built to do so. Within that restriction, the branch is valuable because it introduces a distinction the other branches do not — the distinction between an assay and a surveillance network — and because the public-health community has developed an evaluation discipline from which the other domains can learn.

At the level of the individual biosensor, the governing quantities are familiar cousins of the trade-space already described. The limit of detection is the minimum amount of an analyte the sensor can detect, and the sensor’s usefulness depends on that limit lying below the threshold relevant to its intended use; detection time and specificity are co-equal design constraints alongside it.[28] An assay, in other words, is evaluated much as any detector is: sensitivity, specificity, time to result, and the conditions under which these hold. But the branch’s important move is to distinguish this assay-level evaluation from the evaluation of a surveillance network, where the unit of detection is not an individual assay result at all but an alert generated from aggregated data streams.

This distinction prevents a category error that the framework flags as characteristic: the invalid comparison between a highly sensitive assay and an early-warning network, as though a better assay and a better network were the same kind of improvement. They are not. A multisource analysis of the ESSENCE syndromic-surveillance system tested sensitivity and timeliness at practical false-alert thresholds and reported a one-day median alert lead time among the outbreaks studied, while emphasising that its sensitivity and specificity could still be improved.[29] The relevant evaluation for such a network is not the analytical sensitivity of any assay but the network’s source coverage, completeness, reporting delay, event-level sensitivity, predictive value, and the cost of investigating false alerts — and, above all, the time from exposure through detection to intervention. The public-health framework anchors timeliness to that whole exposure-to-intervention interval, with interim milestones, and warns that the choice and quality of data sources shape the comparison.[30]

The exposure-to-intervention interval that the public-health framework places at the centre of surveillance evaluation is worth dwelling on, because it is the biological branch’s distinctive contribution back to the whole class. In the physical-detection branches, timeliness is typically measured from the appearance of a signal to the delivery of a track — a short interval, seconds to minutes, bounded by the physics of the engagement. In the surveillance branch, timeliness is measured across a far longer and more consequential chain: from the moment of exposure, through the accumulation of a detectable signal in the data streams, through the generation and investigation of an alert, to the initiation of an intervention that changes the outcome. This longer framing carries a lesson the physical branches can absorb: that the decision the detection serves, and the interval within which that decision must be made to matter, are properties of the downstream response and not of the sensor, and that an evaluation which measures only to the point of alert has measured only part of the interval that determines whether the detection was useful. The surveillance branch, by anchoring its timeliness to intervention rather than to alert, models a discipline the whole class would benefit from adopting: measure to the decision, not to the detection.

The branch closes with the most important caution the framework carries, and one with force well beyond the biological domain. A common-data evaluation of biosurveillance detection algorithms under the Bio-ALIRT programme used explicitly bounded false-alert rates and noted plainly that performance against a real biological attack cannot be inferred directly from historic outbreak data alone.[31] This is the framework’s general problem of ground truth, in its sharpest form: where the event of real interest is rare or unprecedented, the data available for evaluation are drawn from a different and more benign distribution, and a system’s performance on the available data may not transfer to the event that matters. The prescription — developed generally in the next section — is to use controlled simulation or injection where real ground truth is scarce, to label the assumed generative model explicitly, and to report transfer uncertainty rather than presenting performance on historic data as if it were performance against the real threat. The biological branch states this caution most clearly, but it binds every branch of the class.

8. A recommended modelling and validation architecture

The framework’s constructive core is a single modelling prescription that unifies the four branches: model the detection system as a directed sensing-to-decision chain, and for each scenario in the matrix estimate the performance of each layer of the chain explicitly. The chain has five estimation points, corresponding to the five layers, and the discipline is to estimate each rather than to collapse them into a single figure.[32][33][34]

At the detection layer, estimate the detection probability or sensitivity, the false-alert rate, the precision or predictive value, and the calibration of these against the decision threshold. At the association layer, estimate the probability of correct association, the false-track rate, the time to track confirmation, the degree of track fragmentation, the rate of identity switches, and the localisation and classification error. At the fusion layer, estimate the incremental value of fusion relative to each sensor used alone, the system’s robustness to a missing or degraded sensor, and — easily forgotten and often decisive — the dependence structure of errors across sensors, since fused sensors whose errors are correlated deliver far less than their independent combination would suggest. At the decision layer, estimate the alert-to-action latency, the probability that a track of sufficient accuracy is delivered inside the decision window, and the false-response workload the system imposes. And at the robustness layer, estimate the degradation curves of all the preceding quantities across the scenario matrix, including the defined adversary model wherever the domain is adaptive.

The architecture then imposes a validation discipline with a specific ordering, and the ordering is the point. Component testing should be used for diagnostic attribution — to locate which layer is responsible for a failure — but the primary performance claim must be reserved for end-to-end trials that inject representative clutter, data delay and loss, and handoff failures. A model should be validated first against measured component characteristics, and only then against independent, held-out, end-to-end exercises; a system that validates well component-by-component may still fail at the seams between components, which is exactly where the hypersonic case showed the real uncertainty to lie. The sequence — components for diagnosis, end-to-end for the claim — is what prevents a collection of individually validated parts from being mistaken for a validated system.

It is worth making the directed-chain estimation concrete, at the level of structure rather than of any specific numbers, because the shape of the calculation is itself part of the framework’s discipline. Consider a single scenario drawn from the matrix — a defined signal-to-clutter condition, target density, geometry, sensor-availability state, and decision-time budget. Within that scenario, each layer contributes a conditional quantity, and the end-to-end performance is a composition of them rather than any one alone. The detection layer contributes a probability of raising a correct alert and a rate of false alerts, both conditioned on the scenario. The association layer contributes, conditional on a correct alert, a probability of forming and holding a correct track, together with the distribution of the time that formation takes. The fusion layer modifies these conditional on which sensors are available and on the correlation between their errors. The decision layer contributes, conditional on a track of adequate quality existing, the probability that it is delivered within the decision window. The composed end-to-end quantity of interest — the probability of a timely, correctly associated, decision-adequate detection — is the product of these conditional terms, integrated over the uncertainty in each, and it is almost always substantially lower than the headline detection probability that a favourable demonstration would report, because each seam multiplies in a further conditional probability below one.

Two features of this composition deserve emphasis because they are the features a single-figure evaluation destroys. The first is that the conditional terms are not independent: a scenario that stresses the sensing layer often stresses the association layer too, so the joint degradation under stress is typically worse than the product of the separately measured degradations would suggest, and an evaluation that measures each layer under favourable conditions for the others will overstate the whole. The second is that the composition is where correlation among sensor errors does its damage: if two fused sensors fail together under the same clutter condition, their combination provides far less resilience than their individual reliabilities imply, and only a model that carries the cross-sensor dependence structure explicitly will price this correctly. The framework’s insistence on estimating each layer separately and then composing them, rather than measuring the whole in one favourable pass, exists precisely to surface these two effects, which a headline figure is structurally unable to show. The numbers in any real instance of this schema are, of course, system-specific and are not supplied here; what the framework supplies is the structure of the calculation and the discipline of refusing to collapse it.

Where ground truth is scarce, as it is for rare physical events and unusual outbreaks alike, the architecture prescribes the discipline the biological branch stated most sharply: use controlled simulations or injections in place of the missing real events, but label the assumed generative model explicitly, and report the resulting transfer uncertainty rather than presenting simulated success as field performance.[35][36] This is the single most important honesty constraint in the whole framework. A simulated result is a statement about a model of the world, and its value depends entirely on how well that model matches the world in the respects that matter. To present a simulation result without labelling its generative assumptions, or without reporting how much uncertainty the simulation-to-reality transfer introduces, is to present a favourable demonstration in its most seductive form — one in which the favourable conditions are not merely selected but constructed. The framework permits simulation, which scarcity often makes unavoidable, but only under the discipline of stated assumptions and reported transfer uncertainty.

9. What the evidence supports, and the inference the author will commit to

The cross-domain evidence assembled here supports one methodological conclusion, stated as strongly as the evidence permits and no more strongly. Detection performance is a property of a system operating under conditions, not a property of an isolated sensor or classifier. Sensor diversity can improve coverage and resilience, but only if fusion is assessed with explicit penalties for false tracks, latency, and cross-sensor correlation. The most decision-relevant output of an honest evaluation is not a single accuracy figure but a family of threshold-and-condition curves: the probability of timely, correctly associated detection versus false-response burden, stratified by environmental and adversarial stress. And the key missing evidence in many detection architectures — across all four domains — is not another nominal accuracy figure but exactly this end-to-end, condition-stratified evaluation.[37][38][39][40]

Beyond what the evidence establishes, this paper commits in print to one inference that goes beyond the current evidence, labelled as such and defended as reasonable rather than proven. The inference is this: as threats across all four domains increasingly incorporate adaptive and evasive design intent — as targets are built, more and more, specifically to defeat the systems that hunt them — the binding constraint on real detection performance will shift decisively from the sensing and classification layers, where most measurement effort is currently spent, to the robustness layer, where least is. Programmes that continue to report demonstration-favourable numbers will systematically overstate their fielded performance, and the size of the overstatement will be largest exactly where the threat adapts most. The single sentence from the adaptive branch — that a nominal receiver-operating-characteristic curve is not an adversarial one — will, on this inference, become the governing constraint of the entire class, and the programmes that thrive will be those that reorganise their evaluation around the adversary rather than around the demonstration.

Epistemic status of the committed inference: Probable, and explicitly beyond the current evidence. It is an extrapolation from the demonstrated fragility of digital classifiers under adaptation[41][42] to the trajectory of the whole threat class, and it depends on the assumption — reasonable but not proven — that adaptive design intent will generalise across domains as this paper argues it will. It is offered as the inference the author is willing to be held to in print, in the spirit of forecasting ahead of the evidence while labelling the distance travelled beyond it. It could be wrong if physical-domain evasion proves far harder to engineer than digital evasion, and the paper flags that possibility as the principal risk to the forecast.

The honest limits of the paper should be restated at its close. This is a targeted evidence scan, not a formal systematic review; it prioritises open methodological, government-analytical, and primary technical sources, and it does not establish the capability of any classified system or claim comprehensive coverage of any programme. The framework yields condition-stratified curves rather than universal thresholds, and deliberately so: the claim that there is no single number adequate to detection performance is not a gap in the framework but its central finding. The adversarial evidence is strongest in the digital domain and its transfer to physical sensing is argued rather than demonstrated. And the ground-truth problem — that the events of real interest are often too rare to furnish adequate evaluation data — is a genuine and, in some domains, unresolved limit, mitigated by disciplined simulation but not removed by it. The paper offers a framework for honest evaluation, not a claim to have completed one.

10. On classification, scope, and review

A closing word on the boundary within which this paper has been written, because the boundary is part of the method. The paper is an open-source methodological treatment. It draws exclusively on published technical, governmental-analytical, and public-health literature, and it makes no claim about the capability, performance, or vulnerability of any classified system. It is confined throughout to the meta-level of detection and evaluation — how to model and assess detection systems — and it deliberately excludes, as its source evidence scan also excluded, vehicle design, countermeasure tactics, pathogen engineering, and acquisition. The biological branch is restricted entirely to sensing and surveillance. No part of the paper is intended to provide, and no part does provide, operational assistance in building, optimising, acquiring, or evading any of the threats it discusses. Its contribution is defensive: to make the evaluation of detection systems more honest, so that those who rely on such systems are less likely to mistake a favourable demonstration for operational reality.

References

  1. Lee E-H, Song T (2017). Multi-sensor track-to-track fusion with target existence in cluttered environments. IET Radar, Sonar & Navigation. https://doi.org/10.1049/IET-RSN.2016.0497
  2. Centers for Disease Control and Prevention. Framework for Evaluating Public Health Surveillance Systems for Early Detection of Outbreaks. MMWR Recommendations and Reports. https://www.cdc.gov/mmwr/preview/mmwrhtml/rr5305a1.htm
  3. Churchill S, Randell C, Gill EW (2006). Data Fusion: Cumulative Effects of Discrete Fusion on Target Detection Probability. 2006 IEEE International Symposium on Geoscience and Remote Sensing (IGARSS). https://doi.org/10.1109/IGARSS.2006.465
  4. Menendez HD (2022). Measuring Machine Learning Robustness in front of Static and Dynamic Adversaries. IEEE International Conference on Tools with Artificial Intelligence (ICTAI). https://doi.org/10.1109/ICTAI56018.2022.00033
  5. Mulhollan Z, Gamarra M, Vodacek A, Hoffman MJ (2022). Essential Properties of a Multimodal Hypersonic Object Detection and Tracking System. Dynamic Data Driven Applications Systems (DDDAS). https://doi.org/10.1007/978-3-031-52670-1_3
  6. Muller MC, Sander T, Mundt C (2023). Development and testing of a multi-spectral and multi-field-of-view imaging system for space-based signature detection and tracking of hypersonic vehicles. SPIE Defense + Commercial Sensing. https://doi.org/10.1117/12.2663480
  7. U.S. Government Accountability Office (GAO-22-105075). Missile Defense: Better Oversight and Coordination Needed for Counter-Hypersonic Development. https://www.gao.gov/assets/gao-22-105075.pdf
  8. Hommes A, Shoykhetbrod A, Noetel D, et al. (2016). Detection of acoustic, electro-optical and RADAR signatures of small unmanned aerial vehicles. SPIE Security + Defence. https://doi.org/10.1117/12.2242180
  9. European Commission, Joint Research Centre. Detection, Tracking, and Identification of Drones: an Overview on Counter-UAS Techniques and Open Challenges. https://publications.jrc.ec.europa.eu/repository/handle/JRC140546
  10. Pereira A, Warwick S, Moutinho A, Suleman A (2024). Infrared and Visible Camera Integration for Detection and Tracking of Small UAVs: Systematic Evaluation. Drones. https://doi.org/10.3390/drones8110650
  11. Rathore H, Samavedhi A, Sahay SK, Sewak M (2022). Are Malware Detection Models Adversarially Robust Against Evasion Attack? IEEE INFOCOM Conference on Computer Communications Workshops. https://doi.org/10.1109/infocomwkshps54753.2022.9798221
  12. Varshney M, Mallikarjunan K (2009). Challenges in Biosensor Development: Detection limit, detection time, and specificity. https://doi.org/10.13031/2013.29328
  13. Burkom H, Elbert Y, Feldman A, Lin JS (2004). Role of data aggregation in biosurveillance detection strategies with applications from ESSENCE. MMWR Supplements. https://doi.org/10.1037/e307182005-014
  14. Siegrist D, Pavlin J (2004). Bio-ALIRT biosurveillance detection algorithm evaluation. MMWR Supplements. https://doi.org/10.1037/e307182005-027