Stress-Extended VV&A: a proposed framework for bounding the credibility of commercial-off-the-shelf simulation in defence systems
Abstract
Defence organisations increasingly rely on commercial-off-the-shelf (COTS) software — including simulation, modelling, and data-service components — in systems and decisions for which those components were never specifically designed or evidenced. This paper argues that the established discipline of verification, validation, and accreditation (VV&A) already supplies the correct language for governing that reliance, but that its standard application stops short of the conditions under which defence systems actually operate: adversarial manipulation, degraded data, contested electromagnetic environments, and edge deployment. We propose Stress-Extended VV&A, a disciplined extension rather than a new certification regime, which treats hostile and degraded operating conditions as declared elements of a component’s intended-use domain and links the resulting evidence to a bounded accreditation decision. We make three claims only: that credibility claims must be bounded to a named function, decision, configuration, and stress envelope; that evidence must be scenario- and change-aware; and that accreditation must expose residual risk rather than conceal it behind a pass/fail result. We are explicit throughout about what the current evidence base does and does not support, and we identify the empirical gaps — coupled stressors, contested-electromagnetic operations at the system level, and repeatable re-accreditation rules for evolving COTS ecosystems — that the framework cannot yet close.
1. Introduction: the assurance problem, stated conditionally
Modelling and simulation is not a peripheral tool in modern defence. It informs training, analysis, acquisition, and mission-critical decision-making, and in many cases it is the only available basis for judgement, because the actual performance of a weapon or system against a live adversary is unavailable and modelled representations must stand in for it.[1] Where simulation carries that weight, the question of whether a given simulation can be trusted for a given decision is not academic. It is the whole of the matter.
The established discipline for answering that question is verification, validation, and accreditation. VV&A has been characterised for decades as the cornerstone of simulation correctness and credibility,[2] and it supplies a structure this paper will lean on heavily. But the contemporary defence environment introduces a complication that the classical framing did not centre: an increasing share of the components carrying decision weight are commercial-off-the-shelf, procured from a commercial market, evidenced for that market, and integrated into high-consequence defence applications for which their original evidence was never intended.
It is tempting to frame this as a story about COTS being intrinsically less reliable or less secure than bespoke defence software. That framing is wrong, and we reject it at the outset. COTS components can reduce cost and development time, and can benefit from the broad use, testing, and maintenance that a large commercial market provides. The problem is not inferiority. The problem is evidence portability: evidence generated for a commercial market does not automatically support a specific, high-consequence mission claim.[3] The safety-critical software literature makes the point precisely — the argument that a component is ‘proven in use’ transfers only to contexts similar to those in which the use was proven.[4] A simulation tool with an impeccable commercial track record in a benign environment has, on that basis alone, established nothing about its behaviour under electronic attack.
This reframing sets the question the paper answers. It is not ‘is COTS safe?’ but a conditional: what additional evidence is necessary for a particular component, in a particular integration, supporting a particular mission function, across a particular operating envelope? We state our thesis as a proposal rather than a settled doctrine. Stress-Extended VV&A adapts intended-use credibility assessment to COTS-enabled systems operating under specified hostile or degraded conditions. It does not, and cannot, ‘certify COTS’ in general.[5]
A note on the evidence base. This paper is built on a literature that is real, peer-reviewed, and directly relevant, but uneven. The VV&A and COTS-assurance foundations are strong though in places older; the adversarial-robustness literature is technically rich but bound to image-classification benchmarks; and direct empirical work connecting COTS behaviour to contested defence environments was sparse in the retrieved material. We have written the argument to the strength of the evidence, and we flag the gaps explicitly where they fall rather than papering over them.
2. The established foundation: VV&A, credibility, and intended use
The paper’s strongest foundation is the one it did not invent. The VV&A literature draws a distinction that does the central analytical work of everything that follows. Verification asks whether a modelling-and-simulation implementation is correct — whether the model was built right. Validation asks whether the implementation represents the real system with satisfactory accuracy for the study objective — whether the right model was built. Accreditation adds a third and distinct act: an official decision that the model is fit, or of sufficient fidelity, for a specific use in a specific context.[6]These are not three names for one activity. They are three separable judgements, and the separation is what allows the framework to see a failure the pass/fail mindset cannot.
That failure is the crux. An artifact can be perfectly well implemented — verified — and can even represent its referent system accurately under nominal conditions — validated — and still be entirely unsuitable for a decision, because its abstraction, its inputs, or its environmental assumptions do not hold in the context where the decision is made. The VV&A framing makes this visible by insisting that credibility is always relative to an intended use, a domain of applicability, and a threshold of satisfactory accuracy, rather than being a general property a tool either possesses or lacks.[7]
The military context sharpens the conditionality. The appropriate VV&A approach is not uniform across models; it varies with abstraction level. Numerical and statistical analysis may suit engineering-level models, whereas mission-level and campaign-level models require data and interface validation, domain expertise, and a larger measure of qualitative judgement.[8] VV&A is, in this reading, a form of risk mitigation — a structured defence against decisions taken on the basis of incorrect modelling-and-simulation outputs.[9] That risk framing is the bridge to everything this paper adds, because stress conditions are precisely where the risk of incorrect outputs concentrates.
VV&A also supplies a lifecycle framing that will matter for the COTS case. Validation, verification, and testing are properly understood not as a final, end-stage gate but as continuous activities aligned with the software development lifecycle.[10] This is the conceptual hook on which the later discussion of vendor updates, dependency changes, and re-testing triggers will hang: if assurance is continuous for bespoke software, it cannot be a one-time event for a COTS component whose vendor ships changes on a commercial cadence the integrator does not control.
Two honest qualifications discipline this foundation. First, VV&A does not automatically solve software assurance, cyber assurance, or operational resilience. Its traditional subject is a model’s relationship to a referent system and an intended study; COTS assurance additionally concerns provenance, latent defects, interface behaviour, patching, supply-chain change, and deployment behaviour. VV&A should therefore be understood as the governance and argument structure for the method — not as a source that already supplies every test method and acceptance criterion.[11] Second, the literature itself cautions against the belief that a single universal military VV&A procedure exists; the field has been described as still in development, with no one recognised set of specific instructions or best practices having emerged.[12] This limitation does not defeat the paper’s contribution. It defines its form: what we propose is a reusable decision process, within which the actual evidence and thresholds are allowed to vary by model abstraction, mission consequence, and threat environment.
2A. Related work and the boundary of the contribution
It is worth situating the contribution against the existing body of work precisely, both to credit what the framework inherits and to mark where it stops. The VV&A literature itself divides into a general strand and a defence-specific strand. The general strand, exemplified by early consolidation efforts, established VV&A as the disciplined basis for simulation correctness and credibility and set out the verification-validation-accreditation triad as separable activities.[13] The defence-specific strand refined that triad for military characteristics — the unavailability of ground-truth performance data, the range of abstraction levels from engineering to campaign models, and the accreditation authority’s role as a designated decision-maker rather than a technical tester.[14] The present paper sits squarely inside this second strand and extends it in one specific direction: the treatment of hostile and degraded conditions as first-class elements of the intended-use domain.
The COTS-assurance literature forms a second, largely separate tributary. Its safety-critical core establishes the evidence-portability problem and the bounded nature of ‘proven in use’ claims;[15] its high-reliability-systems branch supplies the top-down, risk-prioritised integration process;[16] and its defence-integration review synthesises sixty-two sources into a conditional account of when COTS integration succeeds.[17] The military-aviation assurance work contributes the evidence-to-objective mapping and the analysis of tolerable incomplete evidence that this paper treats as its most important conceptual import.[18] What none of these sources does — and what the present paper therefore cannot claim to inherit — is integrate COTS qualification with adversarial-robustness evaluation, contested-electromagnetic operation, and edge resilience under a single defence-accreditation argument. That integration is the paper’s proposal, and it is offered as a structuring hypothesis to be tested, not as a synthesis of settled results.
The third tributary — adversarial robustness, edge resilience, and electromagnetic-environment fidelity — is the most technically active and the least defence-specific. Each supplies a genuine methodological lesson: name the threat model and vary its severity;[19] treat disconnection and recovery as testable rather than assumed;[20][21] and evaluate environmental fidelity before trusting equipment test results obtained within it.[22] But each was developed for a purpose other than COTS defence assurance, and the paper’s use of them is explicitly as transferable method rather than transferable result. Marking that boundary clearly is itself part of the contribution, because the failure to mark it — importing an image-classifier robustness number as though it characterised a command-and-control system — is precisely the category error the framework is designed to prevent.
3. The COTS risk case: inheritance, integration, and lifecycle
If VV&A supplies the argument structure, the COTS-assurance literature supplies the specific risks the structure must address. The strongest and best-documented of these is information asymmetry. In safety-critical applications, the integrator frequently lacks access to the supplier’s development-process evidence, because the component is delivered and treated as a black box. This creates a direct tension with assurance standards that expect process evidence as a condition of acceptance.[23] The integrator is asked to make a high-consequence claim about a component whose internal construction and development rigour are, by the terms of the commercial transaction, opaque.
This asymmetry is what motivates three of the framework’s later moves — the explicit mapping of dependencies, the surfacing of assumptions, and the deliberate identification of evidence gaps. But it must be handled with care, and here the paper draws a distinction that matters. A COTS integrator inherits the component’s assumptions and its change dynamics; the integrator does not thereby inherit a presumption of defects. The correct phrase is an assurance-inheritance problem, and it is a problem of inheriting unknowns, not of assuming failure.[24] The argument is for integration-specific evidence, not for a default suspicion that the supplier’s work is bad.
Integration risk is, moreover, distinct from component quality, and conflating the two is a common error. A component of excellent intrinsic quality can still introduce unacceptable risk through the manner of its integration. The high-reliability COTS literature accordingly calls for a top-down process to identify, prioritise, and mitigate system-level risk, rather than treating the ‘COTS’ label as dispositive in either direction.[25] The defence-focused review of COTS integration reaches a compatible conclusion from a base of sixty-two sources: appropriateness is conditional on multiple enabling and deterring factors, and successful integration depends on the management of integration knowledge and lessons learned rather than on the component in isolation.[26] In high-reliability applications specifically, commercial products may have been designed against less-stringent requirements and shorter lifecycle expectations, which can affect reliability and maintainability once integrated into a system with a very different operational profile.[27]
The military aviation literature contributes the most transferable idea: a framework that explicitly relates evidence to safety objectives, and that treats the tolerability of incomplete evidence as a question to be analysed rather than assumed.[28] This is the single most important conceptual import in the paper, because it converts the black-box problem from a reason to reject into a decision to be argued. The claim that evidence is incomplete does not establish that residual risk is unacceptable; it establishes that the acceptance decision requires an explicit assurance argument.[29]
That principle disciplines the whole COTS section against a tempting overreach. Opacity is not, by itself, a reason to reject a COTS component. Black-box components can be managed through a repertoire of system-level measures: interface constraints, partitioning, operational monitoring, redundancy, contractual disclosure, provenance controls, and targeted testing. The existence of an evidence gap is the beginning of an assurance argument, not the end of a procurement decision.[30] Equally, the paper must resist importing aviation or nuclear assurance practices as if their regulatory environments transferred with them. Aviation certification offers genuinely transferable ideas — traceability, independent verification, configuration management, evidence-to-objective mapping — but its assurance case may not map cleanly onto dynamic combat missions, classified and shifting threat models, or software updated on a rapid commercial cadence.[31] The comparative move is to extract practices, not to claim regulatory equivalence.
4. From nominal qualification to a stress envelope
The paper’s substantive advance is to make hostile and degraded operating conditions explicit elements of the intended-use domain, rather than leaving them as unstated background assumptions. This is not a departure from VV&A; it is a tightening of it. Validation is already defined relative to a domain of applicability, a study objective, and a threshold of satisfactory accuracy.[32] What Stress-Extended VV&A does is insist that, for defence systems, the domain of applicability must include the conditions an adversary will impose — and that a credibility claim silent on those conditions is, for defence purposes, incomplete.
We organise the stress space into four dimensions, each with a different strength of supporting evidence, and — critically — a fifth cross-cutting category that the four alone would miss.
The adversarial-input dimension has the clearest methodological base, and it yields the framework’s sharpest rule. A large benchmark study of adversarial robustness in image classification found that robustness rankings among systems can reverse when the perturbation budget or the number of attack iterations changes; a single test configuration is therefore inadequate as a basis for a comparative robustness claim.[33] The same literature insists that a threat model must specify an adversary’s goals, capabilities, and knowledge.[34] Together these support a firm assurance rule: any robustness claim must name the attack model and the tested severity range, because a robustness result untethered from its attack configuration is not merely incomplete — it can be actively misleading.
This is also the site of the paper’s most important limitation, and we state it plainly. The adversarial-robustness benchmark evidence concerns image classification under well-defined perturbation models.[35] Its transferable content is methodological — evaluate against a named threat model across a range of attack strengths — and nothing more. It emphatically does not follow that image-classifier robustness results predict the behaviour of a COTS command-and-control, logistics, or simulation system under adversarial conditions. The lesson transfers; the numbers do not.
The degraded-edge dimension has direct, though not defence-specific, support. Edge services trade reliability against latency and geographic distribution, and resilience methods developed for centralised cloud settings may not address the combined hardware, software, and network characteristics of edge operation.[36] A subsequent edge-cloud study reports tests in which resource and control-plane management mitigated a range of faults — congestion, cyber-attack, controller disconnection, overload, and link failure.[37] This gives the framework a concrete, evidenced basis for treating disconnection, resource scarcity, and recovery behaviour as testable conditions rather than as contingencies to be assumed away.
The contested-electromagnetic dimension rests on narrower evidence, and the paper is careful about what that evidence establishes. A complex-electromagnetic-environment simulation study treats sufficient environmental fidelity as a prerequisite for meaningful equipment test, and reports a weighted, receiver-focused fidelity method validated in a field test.[38] This supports the premise that EM-environment fidelity can be evaluated. It does not support the much stronger claim the paper deliberately declines to make.
Stated as a research gap: we found support for evaluating the fidelity of an electromagnetic-environment simulation,[39] but limited direct evidence linking that fidelity to system-level COTS assurance outcomes under electronic warfare. Demonstrating that an EM environment is faithfully simulated is not the same as demonstrating that a COTS-dependent system remains valid while operating inside it. The framework can require the test; it cannot yet promise that passing it settles the system-level question.
The fourth dimension is edge deployment as a structural condition rather than merely a degraded one — the fact that a component is running at the tactical edge, disconnected from the sustaining infrastructure a commercial product typically assumes, changes the assurance question independently of any single stressor. But the most important structural point about all four dimensions is that they must not be presented as exhaustive or independent. A realistic stress event couples them: sensor degradation, adversarial manipulation, communications loss, operator workload, compute throttling, and adaptation by the system itself can arrive together and interact. A four-column checklist would conceal exactly these interaction effects. The framework therefore adds a mandatory cross-cutting category of compound scenarios, and requires the risk-ranking process to select a small number of coupled tests in which stressors are deliberately combined.[40] Coverage of the four dimensions individually is necessary; it is not sufficient.
5. The proposed Stress-Extended VV&A methodology
We now state the methodology as a six-step assurance-case workflow. It is framed throughout as a structured argument linking evidence to a decision, in the manner the military aviation contracting literature recommends for relating evidence to objectives and evaluating incomplete evidence.[41] It is a decision process, not a scoring scheme, and the distinction is deliberate.
Step one: declare the claim. State the intended mission function, the specific decision the simulation will inform, the deployment context, and the consequence of error. This follows directly from VV&A’s objective- and domain-specific logic: a claim that does not name its intended use cannot be validated against it.[42] The declaration is the object everything else in the workflow tests.
Step two: map dependencies and assumptions. Identify the COTS component and its interfaces, the data it consumes and produces, the infrastructure it relies on, the operators who use it, and the vendor and update dependencies that will change it over time. This step directly addresses the information asymmetry and integration risks documented above, surfacing the inherited unknowns so they can be argued about rather than silently assumed.[43]
Step three: construct and prioritise stress scenarios. Draw from the four dimensions — adversarial input, degraded data, contested electromagnetic environment, and edge constraint — and, mandatorily, from the compound-scenario category. Prioritise not by ease of testing but by mission consequence, likelihood, detectability, and recoverability, so that the small number of tests actually run are the ones whose outcomes most change the decision.
Step four: generate evidence using test modes appropriate to the claim. The available modes — laboratory testing, fault injection, scenario replay, hardware- and software-in-the-loop testing, simulation, red teaming, and field exercise — are selected to match the declared claim and the prioritised scenarios, not applied uniformly. The edge-resilience literature’s fault-injection results[44] and the adversarial literature’s threat-model discipline[45] both inform this step.
Step five: interpret results against the claim. Assess correctness, robustness, detectability, recoverability, and change-resilience — and do so as distinct dimensions of the argument, not as inputs to a single pass/fail statistic. The robustness benchmark’s central finding is the warning here: a fixed attack setting can mislead, so a single summary number can conceal a claim’s fragility.[46] Interpretation is an argument about fitness for the declared use, not a verdict.
Step six: make a bounded accreditation decision. The output is one of four dispositions — approve for the stated use, approve with stated constraints, defer pending mitigation, or reject the claim — accompanied by an explicit record of monitoring commitments and re-test triggers. This is the step that closes the loop back to the declared claim and makes the residual risk a matter of record rather than of assumption.
Two qualifications keep the methodology honest, and both constrain how far it may be pushed. First, the five interpretive dimensions — correctness, robustness, detectability, recoverability, change-resilience — are a sound organising device, but they are not a validated measurement framework. The paper does not assign numeric scores across them or imply that they are commensurable, because they are not: a system may score well on recovery and still be unacceptable if a single transient failure makes a mission decision irreversibly wrong.[47] Second, the decision criteria are mission-specific, and no retrieved source supports universal thresholds for acceptable adversarial degradation, downtime, or fidelity. The methodology therefore separates two activities that are often conflated: evidence collection, which can be standardised, and risk acceptance, which cannot, and which requires a designated authority, a hazard or mission analysis, and an explicit rationale.[48]
Configuration management deserves its own treatment, because it is where the COTS lifecycle problem bites hardest. Re-testing on every vendor update is impractical; never re-testing produces stale evidence that silently decays as the component and its environment change. Between these failure modes the framework proposes an impact-analysis trigger model: a change to a dependency, an interface, a deployment topology, a data distribution, a threat model, or a mission role prompts a review, which may or may not require re-test. We are candid that the thresholds for these triggers — how much change, of what kind, warrants what depth of re-evaluation — are an unresolved empirical issue that this framework structures but does not settle.[49]
5A. Applying the workflow: an illustrative walk-through
To show the workflow as a working instrument rather than an abstraction, we trace it through an illustrative claim of the kind defence organisations routinely face — without inventing an applied case or asserting undocumented results. Consider a COTS simulation component used to generate synthetic sensor data that feeds an operator-training decision: specifically, a commercial physics-based radar-environment simulator repurposed to train operators in target discrimination. The example is illustrative and the walk-through asserts no empirical findings; it demonstrates only how the six steps structure the assurance question.
At step one, the claim is declared narrowly: the simulator is to be trusted to represent radar returns with sufficient fidelity that operators trained on it will correctly discriminate targets in the live system, under the specific conditions of the training syllabus. Note what the declaration already excludes — it makes no claim about the simulator’s fidelity under jamming, nor about its use for engineering analysis rather than training. The intended-use logic does its work immediately by bounding the claim.[50]
At step two, dependency mapping surfaces the inherited unknowns: the simulator’s internal propagation model and its assumptions about the electromagnetic environment (which the vendor may not disclose), its data interfaces to the training system, the compute infrastructure it assumes, and the vendor’s update cadence. Each is a point at which commercial evidence may not transfer to the training claim.[51] At step three, stress scenarios are constructed and prioritised: a contested-electromagnetic scenario (does the simulator’s fidelity degrade gracefully or misleadingly when the represented environment includes interference?), a degraded-data scenario, and — mandatorily — at least one compound scenario coupling electromagnetic contestation with operator workload, since the training value collapses if realism fails precisely when the operator is most loaded.[52]
At step four, evidence is generated by the modes the claim warrants — comparison of simulator output against measured reference data across the prioritised scenarios, rather than a single nominal comparison. At step five, the results are interpreted across the five dimensions: not ‘did the simulator pass’ but whether its representation remains correct under the stress scenarios, whether divergence from reality is detectable to an instructor, and whether the training remains valid as the vendor updates the underlying model. The electromagnetic-fidelity literature supplies the method for the fidelity comparison[53] while reminding us of its own limit — a faithful environment simulation is a necessary input to the training claim, not a sufficient demonstration of it. At step six, the accreditation authority reaches a bounded decision: perhaps approval for unjammed-environment training only, with a re-test trigger on any vendor update to the propagation model, and an explicit exclusion of contested-environment training pending further evidence.
The walk-through illustrates the framework’s central virtue and its central limit in one pass. Its virtue is that every assumption becomes visible and every exclusion becomes explicit: no one reading the accreditation decision can mistake unjammed-environment approval for a general endorsement. Its limit is the unresolved translation problem — even having run the scenarios, converting a measured fidelity degradation into a judgement about whether operators will mis-train is a mission-level inference the component-level evidence informs but does not determine. That limit is not a defect of the illustration; it is the honest shape of the problem, and it is why the empirical case study proposed in the discussion is the necessary next step rather than an optional extension.
6. Institutional context: M&S, digital engineering, and interoperability
The framework does not operate in a vacuum; it enters an institutional environment that both motivates it and constrains it. Defence modelling and simulation already serves training, analysis, acquisition, and deployment, and decisions across all four may rest on simulated results.[54] Department of Defense digital-engineering work frames data-driven decision-making across the acquisition lifecycle as a desired engineering practice,[55] and acquisition-focused analysis identifies digital transformation as an enabler for acquisition and sustainment amid growing threats and challenges.[56] This institutional momentum makes the paper’s question more relevant, not less: as simulation and digital engineering are woven deeper into acquisition, the credibility of the COTS components inside them carries correspondingly more weight.
Training offers a concrete illustration of bounded fidelity done well. In a study of Navy F/A-18 live-virtual-constructive training, researchers surveyed thirty training professionals on fidelity requirements to guide engineering trade-offs, having first identified aircrew concerns about realism and training quality.[57] The instructive feature is that fidelity was treated as something to be justified relative to a specific learning or operational decision, not as an undifferentiated property a simulator has more or less of. That is exactly the intended-use logic Stress-Extended VV&A generalises: fidelity, like credibility, is a claim about a purpose, not a global attribute.
Two qualifications about the institutional context prevent the paper from arguing from adoption, which would be a category error. First, interoperability is not credibility. NATO standards work supports the claim that common technical and data standards enable interoperability and reuse, and that a standards profile is needed beyond individual agreements to give organisations a coherent view of the applicable standards.[58] But common standards enable exchange; they do not establish shared semantics, adequate model fidelity, common threat assumptions, or compatible failure behaviour. A federated simulation can be perfectly interoperable and still be invalid for a joint decision, and the paper insists on distinguishing technical interoperability from the assurance claim that the federation remains valid for the decision it informs.[59] Second, digital engineering is an institutional aspiration and a means to improve traceability; it is not itself evidence that a particular model or COTS component is trustworthy. Widespread adoption of modelling-and-simulation and digital-engineering policy strengthens the relevance of the question this paper asks, but only the proposed evidence chain can support a bounded decision.[60]
7. Discussion: what the method does and does not solve
A framework of this kind earns trust in proportion to its candour about its own limits, so we state them directly. Stress-Extended VV&A does not make a proprietary component transparent. It does not guarantee security against future adversaries. It does not substitute simulation for operational testing in high-consequence cases. What it does is make visible — and therefore contestable and improvable — the claim, the assumptions, the stress cases, the evidence gaps, the acceptance authority, and the re-test conditions that nominal qualification typically leaves implicit.[61][62]
The strongest case for the method is practical rather than absolute. Its central value is that it prevents nominal test success from being mistaken for a claim about operation under different conditions. Each supporting literature converges on this point from its own direction. The adversarial benchmark shows that even within a single technical domain, results depend materially on configuration and threat model.[63] The COTS-assurance work shows that safety evidence may be incomplete and context-sensitive, and that ‘proven in use’ is a bounded claim.[64] The defence VV&A work shows that the appropriate methods vary with abstraction and mission context.[65] Stress-Extended VV&A is the structure that holds these three findings together and turns them toward a decision.
The key unresolved problem — the one the framework organises but cannot yet solve — is how to translate component-level degradation into a mission-level risk decision. Knowing that a component’s accuracy falls by some amount under a given stress does not, by itself, tell an accreditation authority whether the mission decision the component informs is thereby compromised. Closing that gap is not a matter of further conceptual expansion. It requires empirical work: a next research phase should apply the method to one bounded case — for example, a COTS-enabled simulation data service operating at the tactical edge — and compare its decision quality, cost, and re-test burden against those of conventional nominal qualification. A single well-documented case study of that kind would do more to validate the framework than any amount of additional conceptual development.[66] We identify this as the natural successor to the present paper and decline to pre-empt its findings here.
7A. Limitations and threats to validity
Beyond the specific gaps noted in place, three threats to the validity of the framework as a whole deserve consolidated statement, because a peer-reviewed proposal should make its own vulnerabilities easy to find. The first is evidential: the framework is assembled from literatures of uneven maturity and differing domains, and its coherence as an integrated method is a hypothesis rather than a demonstrated result. The adversarial evidence is benchmark-bound; the edge-resilience evidence is drawn from commercial cloud-edge settings rather than contested military ones; the electromagnetic evidence establishes fidelity evaluation but not system-level resilience. A reader is entitled to regard the integration as promising-but-unproven, and we regard it the same way.[67][68]
The second threat is the absence of validated thresholds. The framework structures a decision but supplies no universal criteria for what degree of adversarial degradation, downtime, or fidelity loss is acceptable, because no retrieved source supports such criteria and we decline to invent them. This is a deliberate limitation with a cost: two accreditation authorities applying the framework to the same component could reach different bounded decisions, because the risk-acceptance judgement is theirs to make and to justify. We regard this as appropriate — risk acceptance in high-consequence contexts should rest with a designated authority, not with a methodology[69] — but it does mean the framework standardises the argument’s structure without standardising its conclusion.
The third threat concerns the re-accreditation problem, which the framework organises but does not close. The impact-analysis trigger model tells an organisation when to consider re-testing, but not how deeply to re-test, nor how to bound the cumulative evidence decay that occurs as a COTS ecosystem evolves through many small changes none of which individually triggers a full re-evaluation. This is the framework’s most significant open problem for real deployment, because it is where the commercial cadence of COTS updates collides most directly with the deliberate pace of defence accreditation, and we flag it as the aspect most in need of the empirical study proposed below.[70]
8. Conclusion: a disciplined extension, not a new regime
The argument of this paper is deliberately modest in form and, we hope, correspondingly firm in substance. We do not propose a new universal certification regime for COTS in defence systems; the evidence does not support one, and claiming otherwise would repeat the error the paper exists to correct. We propose instead a disciplined extension of a discipline that already exists. VV&A already treats credibility as conditional on intended use, domain, and acceptable accuracy.[71] COTS complicates that judgement because evidence, configuration control, lifecycle assumptions, and operational context are only partly inherited by the integrator.[72][73] Stress testing becomes genuinely useful at the point where it turns adversarial, degraded, contested, and edge conditions into declared parts of the operational domain — and where its outputs are tied to a bounded accreditation decision rather than a bare test result.
The framework reduces to three claims, and it is stronger for making only three. Credibility claims must be bounded: a COTS-enabled capability is not simply ‘trusted’, but supported for a named function, decision, configuration, and stress envelope. Evidence must be scenario- and change-aware: nominal tests, supplier evidence, and historical use are insufficient without an explicit argument for their transfer to the intended operational context. And accreditation must expose residual risk: the output is not merely a test result but an explicit decision carrying restrictions, monitoring, mitigation, and re-accreditation triggers.[74]
That narrower position is what the retrieved evidence supports, and it remains appropriately cautious about the large gaps that remain — coupled stressors, contested-electromagnetic operations at the system level, and repeatable re-accreditation rules for evolving COTS ecosystems.[75] Naming those gaps is not a weakness of the framework. It is the framework working as intended: a structure that makes the boundary between what is known and what is assumed a matter of record, so that a defence decision resting on commercial software rests also on an honest account of what that software has, and has not, been shown to do.
References
Full citations for the works referenced by the footnotes above, grouped by theme. All are real, peer-reviewed or institutional sources; DOIs are given where available and must be preserved.
Verification, validation, and accreditation
Kim J, Jeong S, Oh S, Jang Y (2015). “Verification, Validation, and Accreditation (VV&A) Considering Military and Defense Characteristics.” Industrial Engineering and Management Systems. https://doi.org/10.7232/IEMS.2015.14.1.088
Youngblood S, Pace D, Eirich P, et al. (2000). “Simulation Verification, Validation, and Accreditation.” Johns Hopkins APL Technical Digest.
COTS software assurance and integration
Aas A (2007). “Use of COTS Software in Safety-Critical Systems.”
Rispoli K, Dermarderosian A (2008). “Risk Assessment and Mitigation of COTS Integration in High Reliability Systems.”
Hawkins TG, Gravier MJ (2019). “Integrating COTS Technology in Defense Systems.” European Journal of Innovation Management. https://doi.org/10.1108/EJIM-08-2018-0177
Reinhardt D, McDermid J (2012). “Contracting for Assurance of Military Aviation Software Systems.”
Stress, adversarial, edge, and electromagnetic evaluation
Dong Y, Fu Q-A, Yang X, et al. (2020). “Benchmarking Adversarial Robustness on Image Classification.” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://doi.org/10.1109/CVPR42600.2020.00040
Aral A, Brandić I (2018). “Dependency Mining for Service Resilience at the Edge.” IEEE/ACM Symposium on Edge Computing (SEC). https://doi.org/10.1109/SEC.2018.00024
Moura JMF, Hutchison D (2022). “Resilience Enhancement at Edge Cloud Systems.” IEEE Access. https://doi.org/10.1109/ACCESS.2022.3165744
Zheng-qiu H, Zhongfu X, Lihua W, Wentai C (2017). “Fidelity Evaluation of Complex Electromagnetic Environment Based on Similarity Theory.” International Conference on Intelligent Human-Machine Systems and Cybernetics (IHMSC). https://doi.org/10.1109/IHMSC.2017.161
Defence M&S, digital engineering, and acquisition
Baldwin KJ (2018). “JDMS Special Issue: Transforming the Engineering Enterprise — Applications of Digital Engineering and Modular Open Systems Approach.” The Journal of Defense Modeling and Simulation. https://doi.org/10.1177/1548512917751964
Hutchison N, Tao HY, Clifford M, et al. (2022). “Digital Transformation in Acquisition: Using Modeling and Simulation to Advance the State of Practice.” INCOSE International Symposium. https://doi.org/10.1002/iis2.12922
Sherwood S, Neville K, Sonnenfeld NA, et al. (2015). “Fidelity Requirements for Effective Live-Virtual-Constructive Training of Navy F/A-18 Pilots.” Proceedings of the Human Factors and Ergonomics Society Annual Meeting. https://doi.org/10.1177/1541931215591271
Huiskamp W, Igarza J, Voiculet A (2012). “NATO Modelling and Simulation Standards Profile.”
[55]Baldwin KJ (2018). JDMS Special Issue: Digital Engineering and Modular Open Systems Approach. J. Defense Modeling and Simulation. doi:10.1177/1548512917751964
[56]Hutchison N, Tao HY, Clifford M, et al. (2022). Digital Transformation in Acquisition: Using Modeling and Simulation to Advance the State of Practice. INCOSE Int. Symposium. doi:10.1002/iis2.12922
[57]Sherwood S, Neville K, Sonnenfeld NA, et al. (2015). Fidelity Requirements for Effective Live-Virtual-Constructive Training of Navy F/A-18 Pilots. HFES Annual Meeting. doi:10.1177/1541931215591271
[58]Huiskamp W, Igarza J, Voiculet A (2012). NATO Modelling and Simulation Standards Profile.
