Automation at this level stops being about individual instruments and becomes about the whole chain, from a transmitter's installation, through how a loop is tuned, to whether the safety function behind it has actually been proven to work.
Every PID loop is a negotiation between three terms, and a junior engineer who treats them as knobs to nudge until the trend looks tidy will eventually make a loop worse rather than better. Proportional gain is the loop's stiffness: the further the process value sits from setpoint, the harder the controller pushes the output, in direct proportion. Push too gently and the loop never quite gets there — it settles with a permanent offset, sitting a few percent short of setpoint because the small remaining error is all the proportional term needs to hold the valve where it is. Push too hard and the same stiffness that removes offset quickly also overshoots, because the controller keeps correcting past the point where correction was needed, and the loop starts to ring.
Integral action is what finally closes the gap that pure proportional control leaves behind: it keeps adding to the output for as long as any error persists, however small, until the error is driven to zero. The cost is time — the more integral action added, the faster offset disappears but the slower and more oscillatory the recovery from a disturbance becomes, because the controller is now reacting to an accumulated history of error rather than just the present value. Derivative action looks the other way: it reacts to how fast the error is changing, letting it anticipate a fast upset and apply a correction before the error has fully developed. That anticipation is valuable on a slow, quiet loop like a large tank temperature, and actively dangerous on a noisy one like a flow signal, because derivative amplifies whatever noise is riding on the measurement into large, erratic swings at the output.
A valve that hunts — cycling open and shut with no settled position — is nearly always stiction or an oversized valve working at the bottom of its range, not a badly tuned controller. Before touching the tuning constants, watch whether the valve position trace itself is smooth or stepping in small jerks; a jerky position trace against a smooth setpoint trace points straight at the valve, not the loop.
A single loop works well when the process between valve and measurement is simple and the disturbances are slow. Where that is not true, three arrangements extend what feedback alone can do, and each solves a different limitation.
Cascade control nests one loop inside another: the outer loop measures the variable that actually matters — jacket water outlet temperature, say — and instead of driving the valve directly, it sets the setpoint of a faster inner loop, typically on valve position or steam flow, which then drives the valve itself. The reason this works better than a single loop is speed. The inner loop can correct a disturbance — a supply pressure change at the valve, for instance — long before it has had time to show up as a change in the slow outer variable. For a cascade to help rather than hinder, the inner loop must genuinely be the faster of the two; cascading a slow loop inside another slow loop just adds a second source of lag.
Ratio control holds two flows in a fixed proportion to each other rather than at fixed absolute values — combustion air to fuel is the example every engineer meets first. The wild flow (fuel demand, set by load) is measured, multiplied by the desired ratio, and used as the setpoint for the controlled flow (air). Because the controlled variable tracks a measured input rather than waiting for a downstream error to appear, the proportion is held through a load change rather than only recovered after one.
Feed-forward goes a step further: it measures a disturbance directly and adjusts the output before that disturbance has had any chance to create an error at all. A boiler's three-element level control is the clearest shipboard case — steam flow is measured and used to move the feed valve as soon as a load change begins, while the level and feed-flow measurements trim that move afterwards to correct for any mismatch. Feed-forward is only ever a supplement to feedback, never a replacement for it, because it can only compensate for the disturbance it was built to measure; anything else still needs the feedback loop to catch it.
A transmitter is only as trustworthy as the path between the process and the number on the screen, and most disputed readings turn out to be installation problems rather than instrument faults. The signal convention itself is worth understanding rather than memorising: a 4–20 mA loop uses a live zero, so 4 mA represents the bottom of range and 0 mA represents a broken loop — wire, power supply or transmitter — not a genuine zero reading. A system that cannot distinguish "the tank is empty" from "the cable is cut" is not safely designed, which is the whole reason the live zero convention exists.
Resistance thermometers (Pt100: 100 Ω at 0 °C, rising by about 0.385 Ω per °C) are preferred over thermocouples wherever the accuracy justifies the cost, because their output is a stable, repeatable resistance rather than a small millivolt signal referenced to a cold junction. Their accuracy still depends entirely on installation: a thermowell not making metal-to-metal contact with the element, or packed with air instead of a conducting compound, reads slow and low regardless of how good the sensor is. A two-wire hookup also reads high by an amount proportional to cable run, because the lead resistance is added in series with the element and cannot be told apart from it; three- and four-wire circuits exist specifically to cancel that error out.
Level and pressure transmitters carry their own installation traps. A differential-pressure level transmitter is calibrated against an assumed liquid density, so a change of cargo or a temperature-driven density shift moves the true level away from the indicated one even though the instrument itself has drifted nowhere. Impulse lines that trap air in a liquid-filled leg, or condensate in a dry leg, introduce a static offset that a bench calibration will never catch, because the fault sits downstream of the transmitter. The habit worth building early: when a reading looks wrong, ask how it is measured — what the sensing element actually touches, and what lies between it and the process — before assuming the instrument itself is at fault.
Calibration proves the transmitter is honest about the signal it has been given; it says nothing about whether the process condition reaching the sensing element is the same as the process condition you actually want to know about.
Every control valve with a spring-return or fail-safe actuator has been given a direction to move on loss of instrument air or control signal, and that direction is a deliberate design choice, not an accident of whichever actuator happened to be in the stores. The question the designer asks is simple: on total loss of control, which position — open or closed — leaves the plant in the safer condition? Fuel oil, steam and any other supply that can create or sustain a hazard if it keeps flowing unchecked is arranged to fail closed, so that losing control cuts the source rather than leaving it running unmonitored. Cooling water and lubricating oil are the opposite case: losing that flow is the hazard, so their valves are arranged to fail open, keeping the flow going rather than stopping it at the worst possible moment.
The mechanism is usually a spring in the actuator, sized so the stored spring force — not the instrument air or the electrical signal — drives the valve to its safe position the instant supply is lost. That is why an engineer checking a new or modified system tests fail-safe action directly, by isolating the air or signal and watching the valve actually move, rather than trusting the P&ID's arrow. A valve can be wired, piped and labelled correctly and still fail in the wrong direction if the actuator was fitted the wrong way round or the spring range was never checked against the working pressure.
In an examination answer this is one of the cheapest places to lose marks unnecessarily: a candidate who describes the actuator, the positioner and the signal chain in detail but never states which way the valve actually fails, and why that direction is correct for that service, has answered the wrong question. State the direction first, then justify it against the consequence of the failure — that ordering is what the examiner is listening for.
An alarm is a request for a human action, not a record of a condition. That distinction sounds obvious until you look at a real alarm log from an unattended machinery space and find hundreds of activations a watch, most of them re-announcing a condition the operator already knows about or cannot act on at that moment. An alarm system with too many low-value alarms does not become safer for having more of them; it becomes less safe, because the operator's attention — the actual resource being protected — is spent on the noise and is not available for the one alarm that matters. Alarm rationalisation, deciding deliberately which conditions deserve an alarm, at what priority, and removing or suppressing the rest, is why this is treated as a real part of UMS class notation thinking rather than housekeeping done after the fact.
The unmanned machinery space notation exists because taking the watchkeeper physically out of the engine room removes the ordinary human presence that would otherwise notice a slow leak, a rising bearing temperature or an unusual sound long before any single instrument alarms on it. The alarm and monitoring system has to substitute for that presence, which is why its requirements go beyond simply alarming abnormal values: a dead-man alarm confirms a duty engineer is actually responsive rather than incapacitated, and a chain of bridge and engineers' alarms ensures that if the duty engineer does not respond, the call escalates to someone who can, rather than the plant being left to run itself past an unreported fault.
Overriding an alarm is sometimes the correct response — a genuinely spurious alarm on a known faulty sensor should not be left howling — but an override is a decision with consequences, not a way of making a problem quiet. It needs authority from someone senior enough to own that decision, a written record of what was overridden and why, and a time limit that forces the override to be reviewed rather than forgotten. An override with none of those three is functionally identical to disabling the safety function it was protecting.
A safety instrumented function is the complete chain — sensor, logic solver and final element — dedicated to taking one specific plant to a safe state when a defined dangerous condition is reached, and it is deliberately kept separate from the ordinary control loop so that a fault in normal control cannot also disable the protection meant to catch it. A high-high level trip on a settling tank and its associated quick-closing valve is a safety instrumented function; the level transmitter and controller used for everyday level control are not the same instrument, even though they measure the same tank, because the whole point is that one can fail without taking the other down with it.
The uncomfortable fact about any safety function that only acts on a rare event is that it can sit unused, and therefore untested, for a very long time. A shutdown valve's solenoid can seize, its wiring can corrode, or its logic can be left in a bypassed state from the last maintenance job, and none of that will show up anywhere until the one moment the function is actually called upon — the worst possible time to discover a fault. Proof testing exists to close that gap deliberately: at a defined interval the function is exercised end to end, as close to its real trip condition as can safely be arranged, and the result is recorded rather than just noted as "tested, OK" in passing.
The interval matters as much as the fact of testing at all, because the average confidence in the function over time depends on how long it sits between tests — confidence is highest just after a proof test and lowest just before the next one falls due, so a function tested rarely spends most of its life in its least-proven state. That is why proof test intervals are not arbitrary maintenance dates; they are chosen against how demanding and how critical the function is, and shortening the interval is itself a real way of buying down risk, not just a paperwork exercise.
The three worked examples below move from loop tuning arithmetic, through an instrument installation fault, to the reliability maths behind proof testing — each one needs two or more of the ideas above linked together, not a single formula substitution.
A jacket-water temperature control loop uses a proportional-only controller with a proportional band of 25%, working on a transmitter span of 30–80 °C. A rise in engine load requires the control valve to move by 48% of its travel to hold the new heat-rejection rate. (a) Find the steady-state offset this produces, in °C. (b) Find the proportional band that would limit the offset to no more than 3 °C.
Proportional band, PB₁ = 25% Transmitter span = 30–80 °C (50 °C span) Valve output shift for new load, Δoutput = 48% Target maximum offset (part b) = 3 °C
(a) Find the steady-state offset this produces, in °C
(b) Find the proportional band that would limit the offset to no more than 3 °C
Proportional band and gain describe the same setting two different ways.
So convert to gain first.
With proportional-only control the offset needed to hold the valve at its new position is the output shift divided by the gain, expressed as a percentage of span, then converted into the physical units of the loop.
Six degrees is more offset than the engineers can accept.
To find the band that limits it to 3°C, first express the target as a percentage of span, then find the gain and band that would produce that offset for the same valve movement.
AnswerOffset at PB = 25% is 6.0°C; narrowing the band to 12.5% (doubling the gain to 8) limits the offset to the required 3°C.
The trap: narrowing the band removes offset but also pushes the loop closer to instability — halving the band here doubles the gain, so the trade-off against hunting has to be checked as well, not just the offset arithmetic.
A shaft bearing is fitted with a Pt100 (100 Ω at 0 °C, 0.385 Ω/°C) wired to the alarm and monitoring system on a two-wire cable run with a combined lead resistance of 3.85 Ω. The AMS displays 110.0 °C and initiates a bearing shutdown. The second engineer suspects the cable rather than the bearing. What is the bearing's actual temperature, and is the shutdown justified?
Pt100: R₀ = 100 Ω at 0 °C, sensitivity 0.385 Ω/°C AMS indicated temperature (uncompensated 2-wire) = 110.0 °C Combined lead resistance, both cores = 3.85 Ω
What is the bearing's actual temperature, and is the shutdown justified?
Work backwards from the indicated reading to find the resistance the AMS actually measured at its terminals.
That measured resistance includes both lead wires.
Which a plain two-wire hookup cannot tell apart from the element itself. Subtract the known lead resistance to find the true element resistance.
Convert the true element resistance back to a temperature.
AnswerActual bearing temperature is 100.0°C, 10°C below the alarmed value; the shutdown is not justified on this reading — the fault is in the uncompensated two-wire lead resistance, not the bearing. The circuit should be converted to three- or four-wire.
The trap: treating a high reading as proof of a hot bearing without asking how the temperature is measured — a two-wire RTD run reads high by design over any appreciable cable length, and the fix is a wiring correction, not a machinery shutdown.
A fuel-oil quick-closing valve's solenoid pilot has a known dangerous failure rate of λ = 2×10⁻⁶ per hour. It is proof-tested once a year (8760 hours) under the planned maintenance system. (a) Estimate the average probability that the valve would fail to close on demand, PFD_avg, over that interval. (b) The chief engineer proposes testing it every six months (4380 hours) instead. Find the new PFD_avg and state the effect of halving the test interval.
Dangerous failure rate, λ = 2×10⁻⁶ per hour Annual proof-test interval, TI₁ = 8760 h Proposed six-monthly interval, TI₂ = 4380 h PFD_avg ≈ (λ × TI) / 2 (periodic-test approximation)
(a) Estimate the average probability that the valve would fail to close on demand, PFD_avg, over that interval
Find the new PFD_avg and state the effect of halving the test interval
For a component only exercised at proof test.
The average unavailability over the interval is half the failure rate times the test interval — the fault could have occurred at any point in that period, so on average it has sat undetected for half the interval.
Express that as a percentage unavailability.
The more intuitive way to report it in the ship's safety management system.
Repeat with the proposed six-monthly interval &mdash.
Only the test interval has changed, nothing about the hardware itself.
The relationship is linear.
Halving the proof-test interval exactly halves the average probability of failure on demand — testing more often is not just reassurance, it is itself a direct way of buying down risk.
AnswerPFD_avg ≈ 0.876% on the annual test interval, falling to 0.438% on a six-monthly interval — testing twice as often halves the average probability of failing to close on demand.
The trap: assuming a documented pass at the last proof test means the valve is 'safe' until the next one — PFD_avg is an average over the whole interval, so confidence is highest right after a test and lowest right before the next one falls due.
Output = K_p(e + (1/T_i)∫e dt + T_d de/dt)Full PID — proportional, integral, derivative summedProportional band = 100/K_p %Narrower band = higher gain = less offset, closer to instabilityOffset ∝ 1/K_pIntegral action is what drives it to zero, not gain aloneCascade: outer sets inner setpointInner loop must be materially faster than the outerFail closed: fuel, steam — Fail open: cooling, lubricationDirection is chosen from the consequence of losing control4–20 mA, live zero0 mA = broken loop, not a genuine zero readingPt100: 100 Ω at 0 °C, ≈0.385 Ω/°CTwo-wire hookup reads high by the lead resistance; use 3- or 4-wireThree-element boiler levelLevel + steam flow + feed flow, feed-forward trimmed by feedbackDead-man, bridge and engineers' alarmsSubstitute for the watchkeeper's presence in a UMS space; escalate if unansweredPFD_avg ≈ (λ × TI) / 2Average probability of failure on demand between proof tests