Skip to main content

Health Worklist

The Health page is a worklist: everything in your fleet that needs a human, ordered by how fast it goes bad. It covers every hub type — CCU camera hubs and every alarm communicator (Olarm, FSK, RDC, HYYP, Ajax, FinMon, CleverMail and the rest) — in one list, with a rail on the right for area-wide patterns and receiver paths.

The page loads one fleet snapshot for your VCR and refreshes it automatically every 2 minutes; the header shows how many devices are reporting and when the snapshot was taken. Click Refresh to re-fetch immediately.

The sections​

Video Ordered by how fast it goes bad · 0:20

Rows are grouped by urgency. "Offline" is not one state — a device that dropped 10 minutes ago and one that has been dead for a month need different people:

SectionWhat lands hereWho acts
Went offline · last hourAny hub whose connection dropped in the last 60 minutesOperator — call / verify now
Offline 1–12 hDropped earlier todayOperator — follow up
Fail to test (FTT)Quiet 12 h – 7 d — past its check-in window. Overwhelmingly radios and communicators whose only liveness is their periodic test signalOperator — the classic failed-to-test flow
Cameras & tamperCCU sites with camera issues, plus streams with an active tamper condition (darkness, blur, scene shift, obstruction)Operator / field team
Power & batteryOpen trouble signals (AC fail 301, low battery 302 …) and devices reporting a weak battery — tonight's outages, visible this afternoonOperator — proactive calls
Disarmed sitesEvery site currently not away (folded by default — expand for the after-hours sweep). Sites that are disarmed and offline sort first, flagged "fully unmonitored"Operator — evening arm check
Chronic · 7 d and olderDead for a week or more, including hubs flagged faultyTechnician queue, not ops

Chronic is hidden by default so a stale backlog never buries a fresh drop — the count stays visible ("… chronic hidden") and the Hide chronic toggle reveals the full backlog (the list shows the oldest 80; the header carries the true total).

Every row is one line: drop time and age, site name and address, area chip, device-type chip, a plain-language description ("Offline 43 min — was online until 14:09"), a metric where the type has one (battery %, camera count, ISP, CCU firmware version), and the site's arm-state pill. Click a row to open the site on the Sites page.

Filters that compose​

  • Offline age chips — ‹ 1 h / 1–4 h / 4–12 h / 12–24 h / 1–7 d / 7 d +, each with a live fleet-wide count. Click to filter the list to that age band (selecting 7 d + reveals chronic rows even while Hide chronic is on).
  • Issue chips — Offline, Fail to test, Camera, Tamper, Mains fail, Low battery, Disarmed, Faulty, Chronic — with counts; only issues that currently exist are shown.
  • Area chips — the six busiest areas by item count, plus any area you have selected, so an area still filtering the list stays visible even after it drops out of the top six on a refresh. While any area is selected, a Clear areas chip removes the area filter.
  • Search — site name, address, serial, area or hub type.
  • Group by — Urgency (default), Area, or Hub type. Grouping by hub type is how you see, say, every CleverMail or FSK problem together.
  • Sort — urgency, newest, longest offline, or site A–Z.

The rail​

Areas acting up flags areas with something interesting going on, so you can tell a pattern from a coincidence:

  • Offline wave — 3 or more devices in the same area dropped within the last 4 hours (with the type mix, e.g. "Olarm 11 · CCU 3").
  • Power — 3 or more open mains/battery trouble signals in the same area within 6 hours — the load-shedding signature.
  • Storm mode — areas an operator has put into storm mode.

Click a chip to filter the worklist to that area. One root cause beats N tickets.

Signal paths is binary — okay, or loud. Healthy receivers collapse into a single quiet line ("✓ 7 paths healthy — …"). A broken path gets an attention card whose pill says what is actually wrong:

  • DOWN 68 d — a heartbeat-supervised base silent beyond twice its expected interval (the age is how long).
  • NOT ROUTING — the receiver's serials[] is empty, so it acknowledges and drops every signal while looking healthy. This is the born-dead state every freshly created receiver starts in — register the base/accounts in CleverOps.
  • NO TRAFFIC — nothing has ever routed here. The endpoint may be fine; check the vendor or panel side is pointed at us.
  • QUIET 28 d — a webhook path that has gone silent for more than 24 h.
  • ATTENTION — the receiver's supervision state reports unhealthy.

Each card shows how many radios ride that path — one dead base is one incident, not N site faults. Paused or deactivated receivers are counted quietly and never alarm.

tip

At the start of your shift: glance at Areas acting up (is anything regional?), then Signal paths (is a receiver down?), then work the top of the worklist. If a receiver is down, everything behind it is one incident — don't dispatch on individual FTT rows for radios that ride a dead base.

Staged escalation while a hub stays offline​

The standard hub-offline event (E350 "Communication lost") fires immediately and is the baseline alert. On top of that, your VCR can configure a staged escalation ladder so a long outage doesn't rely on a single notification that scrolls past. The ladder is authored in CleverOps under Settings → Control Room Defaults → Supervision & Escalation (a VCR-wide default, optionally overridden per site).

Each stage fires after a configured delay measured from when the hub went offline, and lands in one of three places:

  • Control room — a follow-up push to on-app team members, a Telegram message, or a freshly raised event in your queue.
  • Management queue — a review item in the offload lane rather than the live queue, for outages that don't need an operator right now.
  • Silent — an audit-trail entry only.

So a hub that's been down for 15 minutes might produce a push at 5 minutes and a management review item at 15, without any operator having to babysit it. The ladder stops as soon as the hub recovers (the R350 "Communication restore" event), and each stage fires at most once per outage.

Configured per VCR

What escalates, when, and where is set by your control-room administrator in CleverOps — see Supervision & Escalation. If no ladder is configured, only the single E350 appears.

A companion supervision condition, Expected signal / FTT, covers a different failure: a hub that quietly stops checking in without tripping the offline detector. Where Hub-offline reacts to a presence drop (E350), Expected-signal reacts to the absence of a check-in over a configured interval — the same devices the worklist's Fail to test section surfaces. The two are coordinated so a hub never escalates under both at once. It's authored on the same Supervision & Escalation page.

Failed-to-test: phoning the client​

When your VCR configures the Expected signal / FTT ladder with "Phone the client" stages, each stage drops a task on your operator stack — an event titled "Phone the client — …" carrying a checklist item to record the call. This is the classic failed-to-test contact flow.

You make the call — CleverCommand never auto-dials anyone. After phoning the client:

  1. On the FTT task's checklist item, click Record call attempt.
  2. Pick the outcome — Answered, No answer, Voicemail, or Wrong number — and add an optional note.
  3. Click Save attempt.

Cross-operator accountability. So one person can't sign off every attempt, the 2nd and later attempts must be handled by a different operator than the previous one (and, if your VCR turned it on, a different shift). If you try to record an attempt that breaks the rule, CleverCommand shows a message like "stage 2 … must be handled by a different operator than the previous attempt" and nothing is recorded — hand the task to a colleague. Every recorded attempt is stamped with who handled it for the audit trail.

Audit trail. When the task is raised, an "Escalated to client" entry is written to the event timeline (phone — or, if the ladder emails the client, an emailed entry too), and your recorded outcome adds a "client contacted" entry — a clear, end-to-end record that the client was contacted for this failed-to-test.

Maven / Advance mode

The "Phone the client" checklist item appears under the event card in Advance / Maven mode (the same place operator checklists render). The task event itself appears in the queue for any monitored site.

Hub status and event monitoring​

Hub status directly impacts your event queue. If a normally active site has suddenly gone quiet, check the Health page — the site may be quiet not because nothing is happening, but because its hub is offline and cannot send detections. A site with an offline hub is a blind spot; prioritise restoring connectivity, especially for high-value sites.