Reliability, availability, maintainability
Reliability is how often each part of the system fails, ideally by failure mode, since a machine rarely fails in just one way. Maintainability is how quickly it is restored once it does. Availability follows from both: the share of time the system is able to run. For a single repairable unit, long-run availability works out to MTBF / (MTBF + MTTR), a summary of what happened rather than an input to a model.
A RAM study combines them for the whole system, so design, redundancy, maintenance, spares and crew choices can be tested before anyone commits. The ReliaSim guide RAM analysis covers the definitions and the established tools in more depth.
Block diagram arithmetic, and the RAM packages that simulate
The starting point is the reliability block diagram. Blocks in series all have to work, so their availabilities multiply; blocks in parallel need only one to work. That arithmetic is exact under its assumptions: every block fails and is repaired on its own clock, and a stop anywhere is a stop everywhere, immediately. It carries no rates and no storage. On the five-machine bottling line in ReliaSim’s case study, the per-machine availabilities multiply to 46.7%, while the line’s tracked OEE was 54.3%: storage let neighbors keep running through stops, and machines that were starved or blocked were not wearing toward their next failure.
Commercial RAM packages go well beyond that arithmetic, and they simulate. ReliaSoft describes BlockSim as using exact computations or discrete event simulation, and its throughput analysis gives each block a processing rate and keeps items a block cannot process in backlog. DNV describes Maros as event-driven simulation accounting for equipment reliability, configuration and capacity, and maintenance logistics, and Taro models refinery and petrochemical networks with intermediate storage. Isograph describes Availability Workbench running Monte Carlo simulation with capacity analysis, buffer models, spares and labor.
For reliability programs, fault trees, maintenance and spares optimization, and production availability of process plants, a dedicated RAM package is often the right tool, and we will tell you when it is. ReliaSim vs BlockSim sets out where each approach fits.
Where discrete rate simulation fits
On a high-speed production or packaging line, availability becomes throughput through the line’s dynamics: accumulation conveyors and surge tables filling and draining, machines blocking and starving each other, machines running at different rates, converters turning bottles into cases, and short stops measured in seconds. That is the problem ChiAha’s RAM studies are built around.
Discrete rate simulation models flow as rates and fires an event only when a rate changes: a failure, a repair, a changeover, a buffer reaching full or empty. Between events, buffer levels are computed exactly, and every short stop stays its own interrupt with its own time-to-failure and time-to-repair distribution. Failure rates can be tied to line speed, so the trade-off between running faster and failing more shows up as an optimum rather than a guess. Because the model measures what each failure mode costs the line, fixes are ranked by output recovered: in ReliaSim’s bottling line case study, two failure modes with near-identical direct losses returned 1.21× and 0.73× their loss when each was removed.
The published method uses both levels. Fischel and Lange’s WSC 2020 model of a multi-line food plant put a reliability block diagram inside each rate-affecting unit operation, carrying up to twenty failure modes on each of more than twenty unit operations, and coupled the units with discrete rate flow through the tanks and conveyors between them. The published Fischel & Lange (WSC 2020) model was rebuilt in ReliaSim and independently validated by Tom Lange: within 1% of both the plant’s measured OEE and the original ExtendSim model, running the same one-year simulation 1,200× faster on the same laptop.
| Question | Most direct approach | Why |
|---|---|---|
| Reliability targets, allocation and fault trees | Dedicated RAM package | Block diagrams and fault trees are its core |
| Maintenance strategy, spares, crews and life cycle cost | RAM simulation with maintenance and cost modeling | Tasks, resources and costs are modeled explicitly |
| Production availability of an energy or process facility | RAM simulation built for those assets | Configuration, capacities, logistics and storage in one model |
| Throughput and OEE of a high-speed line with accumulation | Discrete rate line simulation ChiAha | Buffers, blocking, starving and rate changes drive the answer |
| Buffer sizing, and which failure mode to fix first for output | Discrete rate line simulation ChiAha | Capacity sweeps and per-mode removal, measured on output |
Most questions can be answered more than one way; the table lists where each approach is most direct, not where the others stop.
Per-failure-mode data, not averages
A RAM study is only as good as its failure data. Stops are grouped by failure mode, because a jam, a sensor fault and a micro-stop on the same machine have different causes, fixes and statistical shapes. Pool them and the distribution fits none of them, and the model cannot tell you what fixing one is worth.
Each failure mode gets a time-to-failure and a time-to-repair distribution, because MTBF and MTTR are outputs, not inputs. Time to failure is measured from the previous stop of any cause, not the previous stop of the same cause, which would count other causes’ downtime as running time. Fits are tested with Kolmogorov-Smirnov and Anderson-Darling, then given an engineering read: a Weibull shape below 1 on a mature machine is a reason to check the data, and the right tail of repair time often matters more than the typical stop. Downtime data analysis covers the method in detail.
Handled this way and validated against history, a line model can come within 1% of measured OEE. That is a result of the discipline, not a property of any tool.
Data, then a validated model, then decisions
Frame the decision
A production target, a redundancy or accumulation choice, a maintenance strategy, a capital request. We scope the model to answer that decision, not to model everything, and if a dedicated RAM package fits the question better, we say so at this step.
Assemble the data you already have
Historian downtime records, machine rates, buffer capacities, product mix, schedules. Failure and repair times are fitted per failure mode where the data supports it; gaps are flagged, never silently defaulted.
Build at the right level
Block-diagram logic for failure modes inside a unit operation and for undecoupled equipment; discrete rate simulation wherever storage and rates set output.
Validate against history
The model must reproduce what the line actually did, including availability by failure mode, throughput, and blocking and starving, before it is allowed to predict anything. We show you the comparison.
Run the experiments
Fewer failures against faster repair, redundancy against buffer capacity, speed against failure rate: hundreds of scenarios instead of the three you had time for.
Hand over the decision, and the capability
You get the recommendation and the evidence. Studies run on the same engines as our products, so your team can keep the model working, we can train your analysts, or we stay on as modeling capacity.
Reliability engineers and simulation builders
In 1990, Andrew Siprelle created the modeling approach now known as discrete rate simulation, and the team has published at the Winter Simulation Conference since 1995. ChiAha works with consulting partners from Technology Optimization & Management and with our Doctors of Reliability: retired industrial reliability engineers and R&D leaders with 32 to 36 years each in manufacturing, across reliability engineering, maintenance engineering, data analysis, and modeling and simulation. One of them, Tom Lange, co-authored the WSC 2020 paper described above. Together the team has built models for 300+ organizations since 1995. See our wider manufacturing simulation work.