Outcomes & ROI

Which One Machine Would Stop Your Shop?

User Solutions TeamUser Solutions Team
|
8 min read

Ask a plant manager which machine going down would hurt most, and you get an answer immediately. Ask which one would stop work entirely, and the answer takes a week and is usually wrong. The two questions feel identical and are not, because hurt is about load and stopping is about whether an alternative route exists at all.

The data to answer the second question is already in your routings and your utilization report. EDGEBIC by User Solutions does not ship a continuity report, so this is an audit you run rather than a button you press, and it takes an afternoon. This post covers the method, the arithmetic that turns it into a ranked list, the four responses available once you have one, and what the exercise cannot tell you. It sits under the EDGEBIC results guide.

Load and coverage are different properties

Every work center has two independent attributes that matter here.

Load is how many hours are booked on it. This is what the Work Center Utilization report answers, and it is the number people quote when they talk about critical machines. High load means the machine is busy and its problems are felt quickly.

Coverage is whether the work it does can go anywhere else. In routing terms, coverage exists when a step has a true alternative configured, or when the step is routed to a work center group with more than one active member. Coverage does not exist when the step names one machine and nothing else can take it.

The risk that actually stops a plant lives in the corner where load is meaningful and coverage is zero. A busy machine with a pool behind it degrades when it fails. A moderately loaded machine with no substitute anywhere stops the jobs that need it, and it stops them completely, because a routing step with nowhere else to go simply waits.

Most plants have protected the first category well and never audited the second.

The audit

Three passes, and the only judgment call is in the first one.

Pass one, list the uncovered work centers. Walk your routings. For each step, ask whether it has a true alternative on the step, or is routed to a work center group with at least two active members. If neither, the primary work center on that step is single-sourced for that operation. Collect the distinct work centers that come out of this. In a typical shop the list is shorter than people expect and contains at least one entry that surprises everyone.

Pass two, attach the exposure. For each work center on that list, gather three figures:

FigureWhere it comes fromWhy it matters
Booked hours over the horizonWork Center Utilization reportThe size of the work that has nowhere to go
Open jobs whose routing includes itRoutings crossed with open ordersThe number of customers affected, which is usually the better ranking key
Position in the routingThe routing itselfA step near the end strands work you have already paid to make

That third column is the one people leave out and it changes the ranking materially. A single-source operation at step two of eight fails cheaply, because the material is still raw and the downstream capacity can be resold. The same failure at step seven strands work in process that has consumed seven operations of labor and cannot be sold, delivered, or easily rerouted.

Pass three, rank and cut. Sort by open jobs affected, then by routing position. In practice the top three entries account for most of the exposure and everything below about the fifth is not worth a program.

What the numbers look like

Take a shop with two entries at the top of the list.

The first is a mill carrying 400 hours over the quarter and appearing on 34 open jobs. It has a second unit configured as a true alternative on most of those steps. If it fails, the scheduler books the alternative, which carries its own setup time and its own hours per unit rather than the primary's, so the work runs slower and costs more. That is a degradation with a price tag you can compute from the alternate's own figures.

The second is a plating step carrying 260 hours over the same quarter and appearing on 19 open jobs, at step six of a seven-step routing, with no alternative anywhere. If it fails for a week, those 19 jobs stop where they stand, and they stop holding almost all of their value in work in process. Nothing else in the plant can absorb them, and nothing downstream can start.

By hours the mill looks like the bigger problem. By exposure the plating step is far worse, and it is the one nobody had on a list, because it never causes weekly pain.

The four responses, in cost order

Once the list exists, the choices are ordinary and mostly cheap.

Pool what is genuinely interchangeable. Machines that can already do each other's work are better modeled as a work center group than as alternates configured by hand on every step. The group has no capacity of its own, so this is a modeling change rather than an investment, and the scheduler picks a member by the group's selection strategy on each run. The capacity case for pooling is separate and additional.

Configure the specific fallback. Where the substitution applies to one operation rather than a class of work, add a true alternative to that step with its own setup time and hours. The alternate is usually slower. A slower route beats no route, which is the whole argument, and the mechanism is covered in how alternative work centers keep a hot job moving.

Qualify an outside source before you need it. For steps with no in-house substitute, the fallback is a vendor. Qualifying one while nothing is wrong is a different activity, and a different price, from qualifying one during an outage. See subcontracting in production scheduling for the wider trade-off.

Accept it deliberately. Some exposures are not worth removing. Writing down that you have accepted one, with the number of jobs it represents, is a legitimate outcome and is enormously better than not knowing.

What this audit cannot do

It is a method, not a feature. There is no continuity report in the product. You are crossing the routings with the utilization report by hand, and the quality of the answer depends on your routings being accurate about which machine actually runs each step.

It does not price the outage. The audit gives you jobs affected and hours booked. Converting those into a cost figure requires your own numbers for contribution, penalty exposure, and recovery, and those are business inputs no scheduler holds.

Coverage on paper is not coverage in practice. A true alternative configured three years ago against a machine that has since lost its tooling, its program, or the operator who could run it is a line in a table, not a fallback. Test the top two or three by actually running a job on the alternate.

An alternate cannot rescue work already started. Steps with actuals stay locked to the machine they ran on. If the failure happens mid-operation, the alternative helps the jobs that had not started, not the one in the chuck.

It predicts nothing about likelihood. This is an exposure map, not a reliability model. Pairing it with maintenance history is what turns it into a priority list, and that pairing is a human judgment.

Why the list is worth the afternoon

Because the alternative is discovering the same information during the outage, when every option costs several times more and the customer conversations are already happening. A plant that knows its three uncovered operations can spend a small amount of configuration effort and one vendor conversation to convert two of them into degradations. That is not a dramatic improvement and it does not show up on any KPI. It shows up the week something breaks, in the difference between a slower month and a stopped one.

User Solutions has been building finite capacity schedulers since 1991, and the recurring pattern is that the expensive surprises are rarely the machines people worry about. They are the quiet ones that never had a second route. For the recovery mechanics once something does fail, see how EDGEBIC protects delivery after a breakdown.

The takeaway

Load tells you which machine hurts. Coverage tells you which one stops you. Cross your routings against the utilization report, rank the uncovered work centers by open jobs and routing position, and spend the small amount of effort it takes to give the top two somewhere else to go. Bring your routings and a quarter of load to a walkthrough of EDGEBIC and we will build the list with you, including the entry you did not expect.

Cross two things you already have. From the routings, list every work center that appears on an operation with no true alternative and no work center group behind it. From the Work Center Utilization report, read the hours booked on each of those over your planning horizon. A work center that carries significant booked hours and has no substitute is a single point of failure by definition. EDGEBIC by User Solutions does not ship a continuity report, so this is a method rather than a button, but both inputs are on screens you use weekly.

A true alternative is configured on one routing step and names a specific substitute for that operation, with its own setup time and hours. A work center group is a pool of interchangeable machines that any step can be routed to, and the scheduler picks a member by the group's selection strategy. Use a group when the substitution is general, meaning any of these mills can take this class of work. Use a per-step alternative when the substitution is specific to one operation. For coverage purposes either one removes the machine from your single-source list.

No. A true alternative replaces the primary for that step on that job. The scheduler evaluates the candidates and books exactly one machine, so an alternate adds flexibility rather than throughput. If you want two machines running shares of the same operation at once, that is independent parallel processing, which is a different configuration with different arithmetic.

They are related and they are not the same list, which is why plants get surprised. The bottleneck is the machine that limits throughput when everything is running. The continuity risk is the machine whose absence has no route around it, and those two properties do not have to sit on the same asset. A constraint with three interchangeable units in a pool is a throughput problem and a mild continuity problem. A lightly loaded specialist machine that appears on one operation of forty routings, with no alternate anywhere, is a minor throughput factor and a genuine stop. Plants tend to protect the bottleneck because it is visibly painful every week, and to leave the specialist alone because it never causes trouble until the day it does. Run both lists and treat them as separate questions with separate answers.

Because knowing changes four decisions that are cheaper than a second line. It tells you which vendor relationship is worth keeping warm before you need it, so a step override or a routing change is a phone call rather than a search. It tells you where a spares holding is justified, because the exposure is now a number of open jobs rather than a feeling. It tells you which quotes to think twice about, since work routed through a single-source step carries a delivery risk your standard lead time does not price. And it tells your customers something useful, because a plant that can say which operations are single-sourced and what it does about them is answering a question its competitors usually cannot. None of that requires capital, and all of it requires the list.

Expert Q&A: Deep Dive

Q: We know our bottleneck. Is that not the same as knowing our biggest risk?

A: They are related and they are not the same list, which is why plants get surprised. The bottleneck is the machine that limits throughput when everything is running. The continuity risk is the machine whose absence has no route around it, and those two properties do not have to sit on the same asset. A constraint with three interchangeable units in a pool is a throughput problem and a mild continuity problem. A lightly loaded specialist machine that appears on one operation of forty routings, with no alternate anywhere, is a minor throughput factor and a genuine stop. Plants tend to protect the bottleneck because it is visibly painful every week, and to leave the specialist alone because it never causes trouble until the day it does. Run both lists and treat them as separate questions with separate answers.

Q: The audit says our plating step has no alternate, but we do not have a second plating line and cannot buy one. What is the point of knowing?

A: Because knowing changes four decisions that are cheaper than a second line. It tells you which vendor relationship is worth keeping warm before you need it, so a step override or a routing change is a phone call rather than a search. It tells you where a spares holding is justified, because the exposure is now a number of open jobs rather than a feeling. It tells you which quotes to think twice about, since work routed through a single-source step carries a delivery risk your standard lead time does not price. And it tells your customers something useful, because a plant that can say which operations are single-sourced and what it does about them is answering a question its competitors usually cannot. None of that requires capital, and all of it requires the list.

Frequently Asked Questions

Ready to Transform Your Production Scheduling?

User Solutions has been helping manufacturers optimize their production schedules for over 35 years. One-time license, 5-day implementation.

User Solutions Team

User Solutions Team

Manufacturing Software Experts

User Solutions has been developing production planning and scheduling software for manufacturers since 1991. Our team combines 35+ years of manufacturing software expertise with deep industry knowledge to help factories optimize their operations.

Let's Solve Your Challenges Together