App definition and data modelling

This section describes how to structure your monitor inventory with tags, and how to group monitors into applications with their own health rules, SLA targets, notifications and dependency views.

The data model

Mugnsoft uses a simple, three-level model:

flowchart LR M["Monitors
(exec, URL, API, TCP, UDP, ping,
nslookup, DB, system, SNMP, WMI)"] -- "carry" --> T["Tags"] T -- "define membership of" --> A["Applications"] A -- "can nest" --> A2["Child applications"] T -- "scope" --> U["Users, reports,
downtimes, dashboards"] A -- "scope" --> U
  • Monitors are the atomic unit: each check (a browser script, a URL probe, a ping, an SNMP device…) produces a status and performance data.
  • Tags are free-form labels attached to monitors. They are the glue of the platform: reports, downtimes, user visibility and application membership are all resolved through tags.
  • Applications group monitors (via tags) and/or other applications into a business-level object with its own aggregated status, SLA history and notifications.

Tags are the foundation of everything else, so agree on a naming convention before creating monitors at scale. A scheme that works well in practice:

  • one tag per business service (crm, webshop, intranet),
  • one tag per environment (prod, uat, dev),
  • one tag per owning team (team-infra, team-app),
  • optional tags per tier (frontend, backend, db).

A monitor can carry several tags; every scoping feature (apps, reports, downtimes, user visibility) computes its monitor list as all monitors carrying at least one of the selected tags, plus any explicitly selected monitors.

Defining an application

Applications are managed from the App page. Users in the ADMIN group can create, edit or delete applications; USER-group users can manage the applications matching their tags.

An application is defined by:

  • Display name, description, state (on/off) and an optional logo image (uploaded after saving; shown on the app tile and in the dependency views).
  • Membership — at least one of:
    • tags: every monitor of any type carrying one of these tags becomes a member,
    • applications: one or more existing apps become children of this app (see Nested applications).
  • Business rule, weights, SLA targets and notifications, described below.

Business rule: how the app status is computed

The business rule (bRule) decides how the member statuses roll up into one application status:

Rule Aggregated status
worst (default) the most severe status among enabled members
best the least severe status among enabled members
highestPercentage the most common status among members (ties resolved toward the more severe one)
weighted each member’s severity is multiplied by the weight of its monitor type; the highest weighted severity wins

Disabled members are excluded from the computation.

What counts as a member

An application resolves its members from either its tags or its sub-applications — never both:

  • if the application has tags, its status is computed from the monitors those tags match, and its sub-applications are not considered;
  • only if it has no tags does it roll up the statuses of the applications listed in apps.

Set the rule explicitly

An application whose business rule is empty or unrecognised is scored OK, unconditionally, whatever its monitors are doing. Applications created through the interface get worst by default; ones created by import or by self-registration may not. If an application insists it is healthy while its members are not, check its rule before anything else.

The status is stored, not computed on read

Each application recomputes its own status on a one-minute schedule and stores the result. A parent reads that stored value for each of its children, so a failure climbs a nested tree roughly one level per minute: in a three-deep tree it can take about three minutes to surface at the top. This is normal, and it is why the aggregate can briefly disagree with a monitor you are watching live in the dependency views.

Per-type weights

With the weighted rule, each monitor type (API, DB, discovery, exec, nslookup, ping, SNMP device, WMI device, system, TCP, UDP, URL) has a configurable integer weight (default 1). Give a higher weight to the types that really define user-facing availability — for example, weight browser scripts and URLs above pings — so a failed ping cannot outrank a failed transaction.

SLA: availability and performance

An application already has a health status, rolled up from its members by the business rule. An SLA is a separate commitment layered on top of it, made of exactly two things:

Part What it is
Target the percentage that must be met, for example 99.9
Impact condition which application statuses spend that budget, for example CRITICAL+

There is deliberately no third part. In particular there is no “SLA status”: an SLA is met or it is missed, against one target. Grading the resulting percentage back into a severity — the model used before — invented a second severity scale on top of the real one.

Each application carries two independent dimensions, and performance is opt-in:

Dimension Default target Default impact Compliant Spends the budget Excluded
Availability 99.9 ERROR OK, MINOR, MAJOR, CRITICAL ERROR DOWNTIME, no-data
Performance (opt-in) 99.5 MAJOR+ OK, MINOR MAJOR, CRITICAL ERROR, DOWNTIME, no-data
  • The impact condition is a floor. The availability default, ERROR (shown as “ERROR only” in the form), spends the budget on a hard failure and nothing else: a CRITICAL application status is visibly broken but still compliant. Set it to CRITICAL+ and CRITICAL starts spending the budget too, MAJOR still does not. How much degradation costs availability is a policy call, which is why the default asks for the one thing every install agrees on.
  • Performance never counts ERROR. A service that did not answer is an availability failure; charging the same minute to both commitments would leave the performance figure a strictly worse copy of the availability one.
  • An application with no performance SLA reports —, not 100%. Nothing was promised, so nothing is reported. Degradations still colour the bars; they are informational until you define a commitment for them.
  • Both dimensions are retroactive. They read the application status timeline, which every stored month already has, so changing a target or a condition re-scores the whole history rather than starting a new one from today.

The webserver keeps a monthly SLA history (up to 24 months) computed from that timeline: percentage, cumulated impacted duration and episode count per month and per dimension, plus the error budget — how much of the allowance was spent, 0% untouched, 100% exactly exhausted, above 100% the target was missed.

A stored month records the earliest instant it actually measured. A month whose history begins on the 21st is reported over that window only, never averaged across the weeks nobody observed. While the status timeline still reaches back that far, the timeline is authoritative and the stored month is recomputed from it.

Time the webserver itself was not running is treated the same way. The application status is recomputed on the webserver’s cron, so a stretch with no webserver behind it is a stretch nobody measured — it is marked as no-data rather than carrying whatever status the roll-up was holding when the process stopped, and it leaves both dimensions on both sides of the ratio.

SLA exclusions (shown as SLA corrections in the app editor) — date/time windows with a reason, for example an agreed maintenance — leave the measurement on both sides of the ratio: excluded time is removed from the impacted duration and from the measured total. Declared maintenance is therefore never credited as uptime, and an outage inside a maintenance window is not quietly forgiven either. Changing the exclusion list invalidates and rebuilds the affected months.

Notifications and remediation

Applications have the same alerting options as monitors, applied to the aggregated app status:

  • email, Slack, Teams and PagerDuty notifications, each independently triggered on failure and/or on status change,
  • a notify threshold (notifyStatus: critical/major/minor) and a notify after counter to suppress flapping,
  • optional remediation scripts executed on failure and/or on status change, with a configurable timeout.

Nested applications

An application may contain other applications instead of (or in addition to) tags. Child app statuses are aggregated by the parent’s business rule, which lets you model hierarchies such as:

Digital workplace (parent app)
├── Mail (app: tags mail-prod)
├── Intranet (app: tags intranet-prod)
└── Video conferencing (app: tags visio-prod)

Nesting is resolved recursively everywhere applications are used — including downtime scoping, where selecting a parent app puts all monitors of its descendants in maintenance.

Dependency views

Two visualizations render the app model:

  • Application dependency map — an interactive 2D graph of the applications, their child apps and member monitors, colored by current status. Use it to spot which member drags an application down.
  • 3D aerial view (App 3D) — a 3D scene where each application is a slab laid out on a plane, useful as a NOC/wallboard view. The layout is per user: drag slabs to arrange them, save the layout and the camera position, and optionally upload a custom background.

See Dependency views for how to read, search and act in both.

Importing applications in bulk

Applications can be created via import/export operations with the app monitor type. The CSV line format is:

app;displayName;description;state (on|off);applications;tags;businessRule (worst|best|highestPercentage|weighted);WeightApi;WeightDb;WeightDiscovery;WeightExec;WeightNslookup;WeightPing;WeightSnmpdevice;WeightWmidevice;WeightSys;WeightTcp;WeightUdp;WeightUrl;threshCri;threshMaj;threshMin;notifyStatus;notifyAfter;emailOnF;emailOnSC;emailR;slackOnF;slackOnSC;slackChan;slackTok;teamsOnF;teamsOnSC;teamsWH;pdOnF;pdOnSC;pdAPI;ScriptOnF;ScriptOnSC;ScriptAction;ScriptActionT;image

Either applications or tags is required; businessRule defaults to worst.

Modelling guidance

  • One application per business service, named as the business names it — applications are what appear on reports and wallboards.
  • Membership through tags, not monitor lists: new monitors tagged consistently join the right applications automatically.
  • Use nesting sparingly — two levels (service → macro-service) cover most organizations; deeper trees make status roll-ups hard to reason about.
  • Pick worst unless you have a reason not to: it is the most conservative rule. Move to weighted only once weights are agreed with the service owner.
  • Set the availability target and its impact condition to the contractual values, and record maintenance windows as SLA exclusions rather than deleting history. Define a performance SLA only where one was actually agreed — an application with none reports —, which is a truthful statement, while 100% would not be.

See also

Translations