The short answer
Define an event clock before calculating KPIs. Record detection, acknowledgement, useful-function loss, temporary crop protection, technician arrival, diagnosis, parts available, repair start, physical restoration, functional test and return to normal control. State the affected system, greenhouse area, crop stage, lost capacity and consequence. Use a small set of balanced measures: crop-critical availability, response time, time to safe temporary control, time to verified restoration, repeat failure, planned-work compliance, overdue risk backlog and emergency-work share. Review data quality and causes with growers and technicians before setting targets.
Mean time to repair can fall because staff close work orders after a reset, even if the defect returns. Planned-maintenance percentage can rise while crop-critical backlog grows. Pair every speed or volume metric with recurrence, verification and risk so the dashboard cannot reward unsafe behavior.
Do not average unlike assets without context. A ten-minute irrigation valve failure across one empty bay differs from a ten-minute heating loss across a full winter crop. Weighting by affected area, capacity, crop exposure or consequence may support better decisions, but keep the raw event record available. Review the denominator as carefully as the event count. Availability based on calendar hours, scheduled production hours or equipment running hours answers different questions. Publish the chosen exposure and do not change it when a poor month makes the result uncomfortable.
This guide is an operating-control framework, not a substitute for the approved design, crop protection plan, product labels, manufacturer instructions, employment and safety procedures, local law, or advice from qualified growers, engineers, water-treatment specialists, plant-health advisers, and safety professionals.
What the buyer should control

| Control point | Required record or action | Release evidence |
|---|---|---|
| Event identity | Record unique case, asset, system, zone, crop, shift, detection source, symptom, initial alarm, operating mode, weather context, crop protection used and related work or warranty case. | One event history prevents duplicate or conflicting records. |
| Time states | Capture detection, acknowledgement, dispatch, arrival, diagnosis, temporary control, parts, repair, test, release and final closure with one clock. | Definitions distinguish safe state, restored function and administrative closure. |
| Operational impact | State lost capacity, affected area, crop exposure, manual labor, water or energy waste, quality effect, safety restriction and production loss. | Grower and operations review validates the consequence. |
| Cause and action | Separate symptom, failed item, failure mode, contributing conditions, immediate repair, root-cause status and recurrence prevention. | Evidence supports cause classification and required follow-up. |
| Balanced KPIs | Track availability for defined critical function, response, temporary-control time, restoration time, repeat rate, preventive compliance, emergency share and risk backlog. | Formula, scope, exclusions, source and owner are documented. |
| Data quality | Review missing timestamps, automatic versus manual records, clock mismatch, cancelled work, merged events, planned outages and unverified closure. | Metric confidence is reported with the result. |
| Management action | Link material trend to staffing, training, spares, design, supplier support, maintenance interval, operating procedure or capital work. | Each action has expected effect and later verification. |
Give every open item an owner, due date, status, related drawing or package, and effect on cost, time, quality, safety, and performance. “Discussed,” “in progress,” or “by others” is not a closure record.
A practical workflow
1. Agree what downtime means for each function
A system may be degraded, manually supported or completely unavailable. Define the state for heating, ventilation, irrigation, controls and other crop-critical services.
2. Capture timestamps during the event
Use controller history, alarm service, work order and shift log. Do not reconstruct every time from memory at month-end.
3. Confirm restoration at the physical system
Test command, feedback, protection, alarm and crop-side result. Keep the case open for root cause or recurrence work without inflating functional downtime.
4. Review outliers before publishing averages
One long event may reveal a missing spare or access problem. A cluster of short resets may reveal a recurring control fault. Both can disappear inside an average.
5. Use KPIs to choose action
The review should end with a maintenance, training, stock, design or supplier decision. Remove any metric that nobody uses or cannot define consistently.
Use one controlled register and preserve superseded records. The team should be able to reconstruct which instruction, setting, role and asset condition applied when an event was observed, adjusted, maintained, tested, restored and accepted.
Roles at the operating interfaces
Owner and operations manager
Set the crop, safety, production and business priorities; assign authority; approve operating limits; and make sure urgent decisions can be made outside normal hours.
Grower and plant-health lead
Define crop-sensitive conditions, hygiene zones, scouting evidence, water-quality needs, permitted treatments, release criteria and the response to suspected pests or disease.
Maintenance and controls team
Keep assets, sensors, software, backups, alarms, isolations, spares and work records usable. Report degraded functions before they turn into crop or safety events.
Suppliers and local specialists
Provide scope-specific instructions, competent service, replacement parts and technical evidence. Local professionals must control regulated electrical, pressure, chemical, fire and environmental work.
Use evidence before releasing the next step
Before people, water, chemicals, crops or equipment enter a released area, confirm that the approved method is current, the responsible person has checked the work, exceptions are controlled, affected teams have been informed, and the record can be retrieved. If a condition is not met, state what may continue, what remains on hold, who owns the action, and when it will be checked again. This is more useful than a general statement that the greenhouse is ready.
Common failure modes
| Failure | Buyer response |
|---|---|
| Downtime ends when the technician leaves | End it at verified useful function under the agreed definition, not at work-order activity. |
| Planned outages are mixed with failures | Record both, but classify them so planning and reliability can be assessed separately. |
| Teams compete for the lowest repair time | This encourages resets and hidden work. Include repeat failures, verification and risk backlog. |
| Only equipment hours are recorded | Add affected area, crop state, lost capacity and temporary protection to support business priority. |
Buyer decision questions
When did useful function fail, and when was it proven restored? Was the crop protected by redundancy, manual work or temporary equipment? Which timestamps are automatic? Did the problem repeat? What risk remains in backlog? Which KPI changes a staffing, spare, maintenance, design or supplier decision? Could the target encourage premature closure?
Link the answer to the greenhouse preventive maintenance plan, the greenhouse warranty defect service log, the greenhouse spare parts criticality guide, so operating decisions remain connected across the crop, equipment, maintenance and evidence.
Frequently asked questions
Is mean time between failures useful?
It can be, if failure and operating exposure are defined consistently. Low event counts and mixed assets can make it misleading.
Should planned maintenance reduce availability?
Report planned and unplanned loss separately, then consider total service effect for crop and production planning.
How is a repeat failure defined?
Choose a period and link by asset, failure mode or symptom. Keep the rule stable and review suspected common causes.
What if the site has no CMMS?
Begin with a controlled spreadsheet or log using unique IDs and required fields. Consistent definitions matter more than software at the start.
Turn the requirement into a controlled deliverable
Share the asset list, alarm and shift logs, work-order fields, recent failures, crop priorities and current reports. Chengfei Greenhouse can help define evidence for its supplied systems while the owner approves site-wide KPI and work-control rules.
Contact Chengfei Greenhouse
