How to Reduce Factory Downtime Before It Spreads

A line can be producing good parts at 10:14 and standing idle at 10:16 because of one failed module, a damaged connector or an HMI that will not boot. Factory downtime is rarely just a maintenance issue. It becomes a production, procurement and delivery problem as soon as the correct replacement cannot be identified and obtained quickly.

For controls teams, the first objective is not to replace every suspect part. It is to restore a known working condition with the least risk to the process. That means isolating the fault, confirming the exact part number and revision, checking what is physically available, and making a decision based on lead time, compatibility and cost.

Factory downtime starts with the fault boundary

The fastest repairs begin by separating the failed device from the wider symptom. A stopped conveyor, lost network communications or a machine safety trip may point to a PLC, but the cause could be a power supply, field wiring, I/O module, network switch, drive fault or configuration issue.

Start with evidence that will still be useful if the shift changes or the equipment is escalated to an integrator. Record fault codes, status LEDs, screen messages, supply voltage, network indicators and any recent work carried out on the machine. Photograph terminal arrangements and module locations before disconnecting anything. If a replacement is needed, these details reduce the chance of ordering a similar-looking but unsuitable part.

A good fault boundary answers three practical questions: what has failed, what else may have been affected, and what must be retained for the replacement to run correctly? A failed PLC CPU, for example, is not necessarily a complete PLC system failure. The programme, memory card, firmware level, communication settings and connected I/O may all determine whether a substitute can be installed without further work.

Check the full identifier, not just the family

Ordering by product family is a common source of delay. “Siemens S7”, “Allen-Bradley ControlLogix” or “Omron HMI” is useful for an initial search, but it is not enough for procurement. Use the complete manufacturer part number from the label, including suffixes, series or revision information where relevant.

For a drive, confirm voltage class, power rating, control interface and encoder or feedback options. For an HMI, check display size, communication ports and software compatibility. For I/O, confirm whether the module is analogue or digital, input or output, and the required signal range. A 24 V DC digital input module is not interchangeable with an analogue input module simply because it fits the same rack.

Where labels are damaged or unavailable, use the installed bill of materials, electrical drawings, PLC hardware configuration and photographs of the rack. The aim is to give purchasing a precise requirement rather than an incomplete description that creates back-and-forth during a breakdown.

Reduce factory downtime with a spare-parts strategy

The best time to source a critical module is before it fails. That does not mean stocking every automation part on site. Excess stock ties up cash, can become obsolete and may be stored poorly. It means identifying the small number of components that can stop output for hours or days if they are unavailable.

Prioritise spares by operational consequence rather than purchase price. A modestly priced communication card may be more critical than an expensive motor because it can halt several cells at once. Likewise, a legacy CPU with no standard lead time may deserve a tested spare even if it has been reliable for years.

Review these factors when setting spare levels:

  • The production impact if the component fails, including downstream disruption and missed dispatches.
  • The quantity of identical units installed across the site or across similar machines.
  • Typical supplier availability, repair turnaround and the risk of end-of-life status.
  • Whether a replacement requires programming, configuration files, licences or specialist commissioning.
  • Environmental exposure, such as heat, vibration, contamination or unstable power supplies.
This assessment should produce a controlled critical-spares list, not a spreadsheet that is never checked. Attach an exact part number, approved alternatives where they genuinely exist, location, condition, last test date and the person responsible for review. If a spare is removed during a failure, make replenishment part of the close-out process rather than a task that disappears once production restarts.

New, refurbished and surplus stock have different roles

For current equipment with short authorised-channel lead times, new and sealed stock may be the obvious choice. For discontinued hardware or an urgent like-for-like replacement, refurbished stock can be a sensible route when the condition, testing process and warranty terms are clear.

The right choice depends on the application. A safety-related component, a high-risk process and equipment under a specific customer validation requirement may call for a particular sourcing route. For a legacy production cell where the priority is restoring an established system, a tested refurbished unit may prevent a lengthy outage at a lower cost than a redesign.

Surplus stock also matters. Plants often hold unused PLC cards, drives, power supplies and operator panels after upgrades or decommissioning. Those items may be valuable to another operation facing an urgent legacy repair. Selling surplus can release space and budget while putting hard-to-find parts back into the market.

Build a sourcing pack before the emergency call

When a production asset stops, the buyer should not need to begin the technical investigation from scratch. A short sourcing pack for each critical system makes enquiries faster and more accurate. Keep it with maintenance records and make sure it is accessible to the teams that cover out-of-hours failures.

The pack should include the exact part number, manufacturer, revision details, photographs, quantity required and condition preference. Add the machine name, process role, whether programming or parameter transfer is required, and the latest known working backup. If there are acceptable substitutes, record who approved them and what engineering changes are needed.

For multi-brand plants, this is particularly useful. A site may run Siemens PLCs, Allen-Bradley drives, Mitsubishi motion equipment, Schneider power products and Omron sensors or HMIs. A single breakdown can involve more than one control ecosystem, so procurement needs a supplier able to source by part number rather than forcing a broad brand-only search.

An independent industrial parts supplier can be useful where a manufacturer channel has extended lead times, a product is obsolete or a same-day replacement is required. Confirm the supplied condition, availability, testing information, dispatch cut-off and returns process before placing the order. Independent resellers are not necessarily authorised by the original manufacturers, so this should be clear in the purchasing decision and compliance record.

Avoid replacing the fault with a second fault

A replacement arriving quickly is only half the job. Many repeat stoppages happen because a module is installed without checking the condition that caused the original failure. If a power supply failed after repeated brownouts, fitting another unit without investigating the incoming supply may only restart the clock. If an output module failed because of a shorted solenoid, the replacement can fail immediately.

Before energising a replacement, check field devices, fuses, earthing, cabinet temperature, cooling fans, cable damage and power quality as appropriate. Follow the equipment manufacturer’s isolation and safety procedures. For programmable devices, verify that the correct programme, firmware and parameter files are available, then preserve the faulty unit for later inspection where practical.

It is also worth confirming that the replacement is genuinely equivalent before the machine is returned to automatic operation. A matching base part number can still have a different revision, memory specification or communications capability. Run controlled functional checks, monitor fault history and involve the relevant controls engineer if the process has changed.

Treat every urgent repair as data

Each downtime event should improve the next response. Record the failed part, root cause where known, hours lost, replacement source, lead time and whether the spare held on site was sufficient. Over several events, this creates evidence for better stock decisions and highlights equipment that needs planned migration rather than another emergency repair.

The most useful measure is not simply how many spares are in the store. It is how quickly the site can identify a failure, confirm the right part and return the asset to stable production. Keep critical part numbers accurate, keep programme backups current, and know where the difficult legacy items can be sourced before the line stops.