Skip to main content
Business software? Visit Kipeo Digital ↗
HarunLucas.com
Article
Engineering SystemsPublished

Engineering Thinking in Software Development: Designing for Failure, Limits and Reliability

11 min read

Software that runs is not automatically a well-engineered system. This article explores how failure analysis, requirements, limits, interfaces, verification and maintainability can shape more reliable software-intensive engineering systems.

Engineer reviewing software architecture and system data beside physical engineering equipment.

Software can run correctly and still be poorly engineered.

A feature can work in a demonstration, an API can return the expected response, and a dashboard can display the right information—yet the overall system may still be fragile when network conditions change, data becomes invalid, a dependency fails or somebody needs to modify the code six months later.

That distinction became clearer to me while developing software systems myself. I have encountered situations where an implementation technically worked, but continued development revealed the need for better structure, error handling, testing, modularity, validation or reliability.

My mechanical-engineering background also influences how I approach those problems. Instead of seeing software only as lines of code, I tend to think in terms of inputs, outputs, components, interfaces, operating conditions, constraints and what happens when one part of the system stops behaving as expected.

That is where engineering thinking becomes useful. It does not mean treating software as though it were a shaft, bearing or pressure vessel. It means applying the broader discipline of engineering to software-intensive systems.

NASA's systems-engineering framework itself treats modern systems as combinations of hardware, software, people, processes and other interacting elements rather than separating software from the larger engineered system.

Start with requirements, not code

Mechanical design rarely begins with:

“Which material do I want to use today?”

The engineer first needs to understand what the component must do.

  • What loads will it experience?
  • What environment will it operate in?
  • How accurate must it be?
  • What constraints apply?

Software development benefits from the same order of thinking.

A technology-first question might be:

“Should I build this with React, Python or Node.js?”

An engineering question comes earlier:

“What must this system do, under what conditions, and how will we know that it does it correctly?”

Consider a machine-monitoring dashboard. “Display machine information quickly” is a vague requirement.

A more useful requirement would identify:

  • Which machine information must appear
  • How current the information must be
  • What operating conditions apply
  • What should happen when data is unavailable
  • What level of delay is acceptable

Requirements engineering turns an intention into behaviour that can eventually be tested.

The current ISO/IEC/IEEE 12207:2026 standard reflects this broader life-cycle view of software, covering processes across development, operation, maintenance and eventual disposal rather than treating coding as the whole engineering activity.

Ten engineering principles and their software-system counterparts, from requirements and failure modes through to maintainability and trade-offs.

Design for failure, not only normal operation

One of the most transferable engineering habits is asking:

“How can this fail?”

A mechanical engineer considers fatigue, wear, overheating, fracture, misalignment, overload and other failure mechanisms. The specific mechanisms are different in software, but the question remains useful.

A software-intensive engineering system might encounter:

  • A sensor that stops reporting
  • A database that becomes unavailable
  • An API request that times out
  • Malformed machine data
  • A lost network connection
  • Duplicated data
  • Insufficient storage
  • Excessive processing load
  • An unavailable third-party service
  • A model that returns an invalid prediction

Designing only for the expected path creates a system whose reliability depends on everything else behaving perfectly.

Good engineering asks the next question:

“If this component fails, what should the rest of the system do?”

It might:

  • Retry
  • Buffer data
  • Reject invalid input
  • Switch to a fallback state
  • Raise an alarm
  • Continue operating with reduced functionality
  • Stop a particular process safely

The objective is not to pretend failure can be removed completely. It is to make failure behaviour deliberate rather than accidental.

Failure-mode thinking can improve software design

Failure Mode and Effects Analysis is widely used in engineering to examine possible failures and their consequences. A lightweight form of the same reasoning can be useful when designing software-intensive engineering systems.

Consider a machine-condition platform:

ComponentPossible failurePossible effect
Vibration sensorStops transmittingMachine condition becomes unknown
Data connectionIntermittent communicationMissing or delayed measurements
DatabaseWrite operation failsCondition history becomes incomplete
Prediction serviceInvalid outputMaintenance recommendation may be unreliable
DashboardStale dataUser may believe old information is current

The useful question is not merely whether each component works.

Ask:

  • How will the failure be detected?
  • What consequence does it create?
  • Can another component compensate?
  • Should the system continue operating?
  • What information should the user receive?

That shift—from feature thinking to failure thinking—is one of the clearest ways engineering discipline improves software design.

Define an operating envelope

Mechanical tolerances should not be copied literally into software vocabulary. A shaft dimensional tolerance has a specific physical meaning. Software does not have dimensional tolerances in that sense.

But software systems absolutely have operating envelopes.

A system may have limits involving:

  • Response time
  • Concurrent requests
  • Sensor sampling rate
  • Storage
  • Numerical range
  • Network bandwidth
  • Memory
  • Database connections
  • Processing throughput
  • Acceptable data age

The transferable engineering principle is:

“Define the range of conditions within which the system is expected to perform correctly.”

A system that has only been tested with ten records should not automatically be assumed to behave correctly with ten million. A monitoring system tested on a reliable local network should not automatically be assumed to behave identically over intermittent connectivity.

Limits exist whether we specify them or discover them through failure. Engineering tries to specify them first.

Design margin is useful—but it is not a mechanical factor of safety

Mechanical engineering often provides deliberate margin between expected loading and failure. Software-system design can benefit from similar reasoning, but the analogy should not be taken literally.

A structural factor of safety has a specific mathematical and physical meaning. Software capacity margin is different.

Examples might include:

  • Compute capacity above normal peak demand
  • Unused storage capacity
  • Additional connection capacity
  • Network bandwidth headroom
  • Queue limits that prevent uncontrolled growth
  • Redundancy for important services

The common principle is headroom.

If a system normally operates permanently at its absolute capacity, even a small disturbance may push it into failure.

The engineering question becomes:

“How close to the limit are we prepared to operate?”

Interfaces deserve as much attention as components

Many failures occur not because individual components are badly designed but because assumptions at their boundaries do not match.

Mechanical engineering has interfaces everywhere:

  • Shaft and bearing
  • Fastener and joint
  • Pipe and fitting
  • Motor and transmission

Software-intensive systems have them too:

  • Frontend and API
  • API and database
  • Sensor and gateway
  • Gateway and network
  • Application and external service
  • Prediction model and dashboard

Suppose one component sends temperature in degrees Celsius while another assumes Fahrenheit. Both programs could be executing exactly as written. The system would still be wrong.

NASA's systems-engineering guidance explicitly treats interface verification as part of system verification because integrated behaviour depends on more than individual component performance.

This is why interface definitions matter:

  • Data format
  • Units
  • Timing
  • Expected responses
  • Error states
  • Ownership
  • Versioning

An interface is an engineering boundary. Treat it as one.

Instrument software the way we instrument machines

If the condition of a mechanical system matters, engineers measure it.

Depending on the problem, that may involve:

  • Vibration
  • Temperature
  • Pressure
  • Flow
  • Speed
  • Strain

A software system also needs ways to reveal its condition.

Useful software instrumentation can include:

  • Logs
  • Performance metrics
  • Error counts
  • Request latency
  • Traces
  • Resource utilization
  • Health checks
  • Queue depth
  • Data freshness

The analogy is not exact, but the design principle transfers well:

“A system that matters should provide enough information to understand its operating condition.”

Otherwise an engineer may only discover that something is wrong when the final output fails. Observability turns internal behaviour into evidence.

Verification is not the same as validation

This distinction is fundamental. NASA describes verification as establishing compliance with requirements, while validation establishes that the system meets the intended user or mission need.

Consider a predictive-maintenance application. Suppose the specification says the system must:

  • Receive vibration data
  • Calculate selected features
  • Execute a prediction model
  • Display the resulting condition classification

Testing may confirm that each function works correctly. That is part of verification.

But a harder question remains:

“Does the system actually identify useful equipment deterioration reliably and early enough to support maintenance?”

That is a validation question. A program can satisfy its software specification and still fail to solve the engineering problem it was created for.

This is why successful execution is not the final measure of an engineering system.

Quality is more than functional correctness

A system may return the correct answer but still be unsuitable in practice. Imagine software that produces correct calculations but:

  • Crashes regularly
  • Is extremely slow under realistic load
  • Cannot recover from communication failure
  • Is difficult to secure
  • Cannot be modified without breaking unrelated components

Functionality alone does not describe its quality.

ISO/IEC 25010:2023 formalizes software and ICT product quality through multiple characteristics that are meant to be specified, measured and evaluated.

The Software Engineering Institute likewise treats qualities such as performance, availability, security, interoperability and modifiability as architecture-level concerns.

That leads to an important engineering principle:

“A system can be functionally correct and still be badly engineered.”

Design for maintenance, not just initial completion

Mechanical systems need maintenance. Software does too, although the mechanism is different.

Software does not physically wear like a bearing, but its environment changes:

  • Requirements evolve
  • Dependencies change
  • APIs change
  • Vulnerabilities are discovered
  • Data volumes grow
  • New hardware is introduced
  • New developers need to understand the system

A system that works today but is extremely difficult to modify may create substantial future risk.

Software maintainability therefore benefits from:

  • Modular components
  • Clear responsibilities
  • Stable interfaces
  • Meaningful naming
  • Documentation
  • Automated tests
  • Controlled dependencies
  • Understandable configuration

This is similar in spirit to designing mechanical equipment with access for inspection and service. Maintenance should not first become a consideration after the system becomes difficult to maintain.

Every design contains trade-offs

Engineering rarely provides a solution that maximizes every desirable characteristic simultaneously. Increasing one property often affects another. Software architectures behave the same way.

Increasing redundancy may improve availability but increase cost and operational complexity. Caching may improve performance but introduce data-consistency problems. More security controls may increase protection while also affecting usability or performance. Greater abstraction may improve modifiability but add conceptual complexity.

The Software Engineering Institute's Architecture Tradeoff Analysis Method exists specifically to reason about architecture in terms of interacting quality attributes and the compromises among them.

Engineering thinking therefore does not ask:

“What is the best architecture?”

It asks:

“Best for which requirements, under which constraints, and at what cost?”

Example: engineering a CNC predictive-maintenance system

Consider an AI-based condition-monitoring system for a CNC machine. At first glance, the project may appear to be a software problem:

“Train a machine-learning model to predict equipment condition.”

But the real engineering system is much larger.

CNC machine → sensors → data acquisition → communications → storage → feature processing → prediction model → dashboard → maintenance decision

A CNC predictive-maintenance system spans physical machinery, sensing, edge acquisition, connectivity, data management, software and a human maintenance decision — reliability depends on every stage and the interfaces between them.

Now apply the engineering concepts.

Requirement

  • How current must the condition information be?
  • How accurately must abnormal states be identified?

Failure mode

What happens if a sensor stops reporting?

Interface

How are sensor values transferred into the processing pipeline, and are units and timestamps consistent?

Operating envelope

How much data can the pipeline process reliably?

Margin

How much additional load can the system accept before latency becomes unacceptable?

Observability

How do we know whether the data pipeline and prediction service themselves are functioning correctly?

Verification

Does each software component perform according to specification?

Validation

Does the overall system actually improve maintenance decisions?

Maintainability

Can a sensor, model or database component later be changed without destabilizing the complete system?

This is why I increasingly see software development through an engineering-systems lens. The code is important. But the code is only one component.

My ongoing AI-based predictive-maintenance project similarly combines condition-monitoring data, Python-based analysis and machine-learning experimentation, with future work intended to explore live CNC integration.

Where the mechanical analogy stops

Cross-disciplinary analogies are useful only when their limits are acknowledged.

Software does not physically wear

A bearing can fatigue or wear through physical operation. Software instructions do not fatigue because they have executed many times. Software systems can become unreliable because their environment, dependencies, requirements, data or threat conditions change—but this is different from mechanical wear.

Software limits are not dimensional tolerances

Latency limits or numerical ranges may play a conceptually similar role in defining acceptable operation, but they are not mechanical tolerances.

Capacity margin is not structural factor of safety

Both involve designing away from an unacceptable limit. Their physical meaning and calculation are different.

Software failure mechanisms are different

A mechanical component may fail because of material defects, fatigue or random physical variation. Software faults generally originate in requirements, design, implementation, interactions, data or changing operating environments.

The goal is therefore not to force mechanical terminology into software. It is to transfer the discipline behind the terminology.

What this means for engineering education

Programming education for engineers should go beyond syntax.

Learning:

  • Variables
  • Loops
  • Functions
  • Classes
  • Frameworks

is necessary.

But engineering students also need to learn to ask:

  • What system am I building?
  • What are its requirements?
  • What assumptions am I making?
  • What can fail?
  • What are the limits?
  • What interfaces exist?
  • How will I test it?
  • How will I know it is healthy?
  • How will somebody maintain it later?
  • Does it solve the original engineering problem?

That is where programming begins to become engineering systems development.

The objective is not merely to teach engineers how to code. It is to teach them how to use software as an engineered component inside a larger technical system.

Key takeaway

Software that produces the expected output has passed an important test. But engineering asks more.

  • What happens when a dependency fails?
  • Where are the limits?
  • How do components interact?
  • How can system condition be observed?
  • How do we know the requirements were satisfied?
  • Does the system solve the actual problem?
  • Can it be maintained?
  • What trade-offs were made?

Those questions do not belong exclusively to mechanical engineering, systems engineering or software engineering. They belong to the broader discipline of building systems that are expected to work reliably under real conditions.

“Working code is the beginning. A well-engineered system explains what happens beyond the happy path.”

References and further reading

  • NASA Systems Engineering Handbook, Rev. 2. Systems-engineering guidance covering requirements, interfaces, verification, validation and life-cycle engineering.
  • ISO/IEC/IEEE 12207:2026 — Systems and software engineering — Software life cycle processes. Current common framework for software life-cycle processes.
  • ISO/IEC 25010:2023 — Systems and software Quality Requirements and Evaluation (SQuaRE) — Product quality model. Defines a quality model for ICT and software products.
  • Carnegie Mellon Software Engineering Institute — Reasoning About Software Quality Attributes. Covers software quality concerns including performance, availability, security and modifiability.
  • Carnegie Mellon Software Engineering Institute — Architecture Tradeoff Analysis Method. Framework for reasoning about architecture risks and trade-offs among quality attributes.
02Frequently Asked Questions

A few common questions

It means treating software as an engineered system with defined requirements, operating conditions, interfaces, failure modes, quality attributes, verification needs and maintenance requirements rather than focusing only on implementing features.

Some underlying reasoning transfers well, particularly failure analysis, requirements, operating limits, interfaces, verification, maintainability and trade-off analysis. The physical concepts themselves should not always be transferred literally because software and mechanical systems fail in different ways.

Verification asks whether the system conforms to its specified requirements. Validation asks whether the resulting system actually satisfies its intended use or user need.

Examples include network loss, unavailable databases, malformed inputs, service timeouts, missing sensor data, storage exhaustion and failures at external interfaces. The relevant failure modes depend on the architecture and operating environment.

Not in the same sense as a mechanical structure. Software-intensive systems can still use deliberate capacity or resilience margins—for example spare processing, storage or network capacity—but these should not be confused with mechanical structural factors of safety.

Because software systems continue to change after initial implementation. Requirements, dependencies, security conditions, interfaces and users evolve, so architecture and code should support safe modification and diagnosis throughout the software life cycle.

04Related Insights

More from this archive

05About the Author
Harun Lucas working at his desk, reviewing code and systems dashboards across multiple monitors

Harun Lucas

Mechanical Engineer · Technology Education Researcher · Engineering Systems Developer

Harun writes from the same practice covered on this site — mechanical engineering, technology education research, and engineering systems development — connecting hands-on work with the ideas behind it.

More About Harun