System/software resilience
Socio-technical factors in industrial software engineering
Increasing interconnectivity is a characteristic feature of evolutionary system landscapes,
where inherent system complexity contributes to increased vulnerability to disturbances.
Therefore, adapting to disturbances is of particular interest.
Resilience engineering has established itself as an interdisciplinary engineering paradigm since the early 2000s,
building on insights from systems theory and security research.
One challenge lies in its conceptual approach to risk.
Risk assessment quantifies the probability and consequences of an event to identify critical system components
vulnerable to specific threats.
System resilience, on the other hand, takes a "threat-agnostic" approach.
It is understood as a system's ability to adapt to changing conditions, detect disturbances, and absorb them.
The resilience concept becomes particularly relevant when risks are not sufficiently predictable.
Therefore, measuring resilience requires novel analytical approaches that complement existing risk approaches while
still allowing for differentiation.
Delay -/Holding Attack
In industrial systems, disruptions in messaging systems are considered particularly critical,
as they are essential for the communication of cyber-physical systems (CPS) and
often form a bridge (coupling) between the digital and physical worlds
(for example, in measurement and control technology for motors, valves, and turbines).
In an industrial context, not only availability but also determinism is crucial—that is,
the certainty that a message (I/O signals) arrives within fixed and often very short time windows,
since even slight delays can have significant potential for damage. A delay attack,
also known as a latency or holding attack,
increases the latency of messages by a constant or variable amount.
This exploits protocol weaknesses, for example, in messaging protocols such as
MQTT (CVE-2026-22535, CVE-2026-40046) or OPC UA (CVE-2025-11043),
or vulnerabilities in software systems.
Automated, adaptive systems for responding to security incidents are among the most effective measures
against threat scenarios. This includes integrated automation in “Cyber Threat Intelligence (CTI) frameworks”
to enable rapid containment and remediation of differentiated disruptions.
One example is the STIX/TAXII framework, which can be used to automate the exchange of information between
CTI organizations, security platforms, and government agencies.
STIX stands for Structured Threat Information Expression and can be understood as a standardized CTI language used to describe cyber threats,
attack patterns, campaigns, indicators, malware, threat actors, and more.
TAXII is a protocol for exchanging CTI information and
has been available in version 2.1 since June 2021.
TAXII is used by various CTI platforms, Security Information and Event Management (SIEM) systems, and
Computer Emergency Response Teams (CERTs).
Link(s):STIX Version 2.1 Errata 01, TAXII 2.1
The concept of system resilience (ISO/IEC 9837:2026) consists of three essential capabilities:
(1) detection and assessment of disturbances (t0, see figure),
(2) adaptation/response (t1, t2), and
(3) restoration of consistent system states and data (t3–t5).
To avoid cascading effects (propagation), the highest possible transient stability must be achieved.
A transiently stable state is understood as a consistent state of
elementary system functions (core functionality) that is maintained until
a steady state (above the curve) can be reached again.
Figure: System Resilience (own illustration)
In modern software architecture, system resilience (ISO/IEC 9837:2026) is not understood as “the absence of errors”,
but rather as the ability of a system to maintain its elementary system functions (core functionality) under particular stress or in the event of system malfunctions
(Hollnagel, E., 2011).
This is supported by:
(1) the avoidance of error states,
(2) the response to fault conditions and the maintenance of elementary system functions, and
(3) the restoration of the steady state (target state).
The target state is considered a stationary state in which the system has the required performance (Above the Curve).
Source: Hollnagel, E. (2011) „Resilience Engineering in Practice: A Guidebook“
Resilience engineering
An expanded scope of tasks in software engineering
Resilience engineering is a relatively new discipline that expands the spectrum by focusing on resilience against
random (unmotivated) and provoked (motivated) influences that can potentially lead to system failure
(“drift into failure”).
A fundamental observation that underlies resilience engineering is that complex systems are dynamic and therefore unstable.
An unstable state can “evolve” or appear “unexpectedly”.
describes in „Programs, life cycles, and laws of software evolution“ how
productive software systems are subject to constant pressure to change and become increasingly complex and
difficult to manage unless measures are taken to reduce this.
The principle remains as relevant today as ever. Especially under changing conditions,
such as the increasing networking of previously isolated and predominantly internal organizational infrastructures,
the field of resilience engineering is in ever-greater demand.
This is particularly true for information systems that process sensitive operational, manufacturing, or research data,
as systems with high coupling and low transparency are especially susceptible to cascading errors and propagation effects.
Source: Lehman, M. M. (1980) „Programs, life cycles, and laws of software evolution“
Socio-technical methods
A multi-perspective view
Socio-technical methods in software engineering consider individual, organizational (social),
and technical factors. Applying these methods can contribute to the design of effective and
efficient organizational structures, methods, and processes.
The development of sociotechnical methods in software engineering stems from the realization that
an isolated consideration of the task spectrum fails to do justice to its true complexity and
existing interrelationships.
For example, interactions can be described between the levels of
(1) individual,
(2) interaction/situation, and
(3) organizational/structural framework conditions,
which are distinguishable globally and locally and change over time.
Global influences are imposed on the organization, the situation, and the individual from the outside,
while the individual's position and perspective change through interaction within the organization.
This opens up a research approach based on the application of sociotechnical methods and measures.
Socio-technical approaches are rarely considered in organizational contexts,
even though their importance has been recognized and scientifically grounded since the early 1960s.
Empirical studies from recent years suggest that the reasons for their reluctance
to use them lie in the difficulty of practical implementation and a lack of connection between these methods,
technical issues, and internal priorities at the operational or tactical level.
Quality for risk mitigation
The discussion of system vulnerability dates back to 1968.
Organized by the NATO Science Committee, the conference in Garmisch in 1968 addressed the question of
how modern, increasingly computer-based command and control systems could remain functional
under conditions of disruption and overload.
Early analyses highlighted that software systems cannot be adequately protected simply by
being error-free as defined by the software specification; rather, robust structures, redundancy,
and adaptive responsiveness are also required.
The fundamental problems of complexity, interdependence, and the risk of malfunction due
to disruptions were particularly emphasized. The 1968 conference is thus considered a
significant cornerstone of software quality.
In parallel, further areas of development emerged with the goal of increasing the stability of software systems.
This was motivated by the realization that both the development and later the operation of complex systems
require systematic methods and processes.
Friedrich L. Bauer was one of the key figures at the NATO conference.
For him, the term “software engineering” was a conscious choice to counteract the “software crisis” of the 1960s.
Friedrich L. Bauer shaped the conference through his conviction that software development
had to be transformed from a “tinkering stage” into a clean, engineering-based process.
Friedrich L. Bauer (1924-2015) was Professor of Mathematics and Computer Science at the Technical University of Munich (TUM).
Sources: Naur, Peter & Randell, Brian (1968) „Software Engineering: Report of a conference sponsored by the NATO Science Committee, Garmisch, Germany, 7th-11th October 1968“; Naur, Peter (1960) „Report on the Algorithmic Language ALGOL 60. Communications of the ACM, 3(5), 299-314.“
In order to optimally design this website for you and to be able to continuously improve it, so-called cookies are used. Below you can allow all cookies by accepting them or make an individual decision.
Accept cookies | Privacy