System, process, log & folder monitoring

The Sentinel Agent (Discovery Agent) watches a host from four angles: the whole system, individual processes, log files, and directories. Each is defined as a set of rules with severity thresholds, evaluated on a schedule, and turned into a Minor / Major / Critical status.

All four monitor kinds are stored as JSON in the agent and run as independent cron jobs. You define them two ways:

  • Sentinel Agent Settings page — edit the monitors interactively in the agent’s Web UI.
  • Component Config Builder — the guided wizard’s Custom Sentinel Agent flow walks you through log monitors, folder monitors, and the rest, then writes a ready-to-run discovery_agent.json.

When a rule breaches its threshold the agent raises an alert at the rule’s severity and routes it through the configured notification channels.

The four monitor kinds

Kind Watches Defined by
Global system Host-wide metrics: CPU, memory, load, swap, disk, network, sockets Per-metric Minor/Major/Critical thresholds
Process Named/discovered processes: CPU, memory, threads, FDs, traffic Per-process thresholds
Log file One file or glob of files: pattern matches, staleness, throughput A path + a list of rules
Directory One folder: file count, total size A path + a list of rules

Schedules use the labels 1MIN, 5MIN, 10MIN, 15MIN, 30MIN, 60MIN. Each monitor runs as an independent job; if a scan takes longer than its interval, the overlapping tick is skipped rather than run concurrently, so a slow disk or a large directory can never stack up overlapping scans.

Global system monitoring

The agent samples host-wide metrics and compares each to its thresholds. Covered metrics include CPU, memory, load average (1/5/15 min), swap (in/out rate and usage), filesystem usage and inodes, disk I/O (read/write rate, queue, utilisation), network traffic in/out, TCP sockets, CPU I/O wait, and database connections.

Each metric has Minor / Major / Critical thresholds, set either as fixed values or computed automatically from a historical baseline (mean + standard deviation). Network traffic thresholds additionally carry an operator (> or <).

For the full field reference, see System Metric Thresholds and Threshold Configuration.

Process monitoring

Processes are monitored either by manual declaration (a process list) or by port scan (the process bound to a port). For each monitored process the agent tracks CPU, memory, thread count, file-descriptor count, and network traffic, each with its own severity thresholds.

For the full field reference, see Per-Process Metric Thresholds.

Traffic needs packet capture

Per-process Traffic In / Traffic Out are the only metrics here derived from captured packets. On a host where the agent cannot capture packets (no Npcap runtime on Windows, or no capture privilege) they stay at 0 and their thresholds can never fire — the Web UI marks those two pills and their charts as not measured rather than idle. Every other metric on this page comes from OS counters and is unaffected. See Bandwidth is always zero.

Log file monitoring

A log monitor watches one file — or a glob pattern such as /var/log/app-*.log — and evaluates a list of rules against the new lines appended since the last tick. The agent tracks a byte offset per file, so each scan reads only what was added; glob patterns pick up newly created files automatically.

Log rotation is handled automatically. Whether the file is rotated by rename-and-recreate (a new file takes the old name) or by copy-and-truncate (the file is emptied in place), the agent detects it and resumes from the start of the new file, so lines are neither silently skipped nor double-counted. Byte offsets are persisted, so a scan also resumes cleanly across an agent restart.

Defining a log monitor

Field Description
path File path, or a glob pattern (*, ?, […]) for multiple files
schedule Scan interval (1MIN … 60MIN)
alertOnMissing on (default) alerts when the file is absent; off stays silent
rules List of pattern rules (below)

Each rule matches lines and alerts on how many matched:

Field Description
name Rule label (shown in the alert)
pattern Regular expression to match on each line
excludePattern Lines also matching this are skipped
operator >, >=, <, <=, =, or first
threshold Match count the operator compares against
severity minor, major, or critical
captureLines Attach the first N matching lines to the alert text as samples (0 = none)
joinLines Group lines into blocks of N and match the pattern against the whole block (0/1 = line by line)
enabled off disables the rule without deleting it

The first operator fires on the first occurrence regardless of count — useful for “alert the moment this ever appears” patterns like FATAL or panic.

captureLines — what lands in the alert

captureLines changes only what the alert shows, never what it counts. The rule always counts every match; capture keeps the first N of them, in file order, and appends them to the alert’s root cause after a | samples: marker, separated by ;:

log /var/log/app.log rule 'errors': 47 occurrence(s) of 'ERROR' > 10 | samples: ERROR db timeout on tenant 12 ; ERROR db timeout on tenant 44

Points worth knowing:

  • First N, not last N. Matches 1…N are kept and the rest are discarded as the scan proceeds, so you get the oldest matches in the tick, not the most recent.
  • Samples only appear when the rule fires. A rule that stays OK produces no root cause, so nothing is shown.
  • One line per sample. When joinLines is also set, only the first line of each matching block is captured — the sample is a pointer to the entry, not the entry itself.
  • Keep N small (2–5). The captured lines go into the alert body verbatim, and log lines can be long.

joinLines — how multi-line entries are matched

joinLines is not a “lines of context before the match” setting. It changes the unit that the regex is applied to: instead of testing one line at a time, the agent fills a buffer with N consecutive lines, joins them with newlines, and tests the pattern against that whole block. A block that matches counts as one match.

With joinLines = 3, a stack trace like this is evaluated as two blocks:

Exception: boom       ┐
  at foo              ├─ block 1 → tested as "Exception: boom\n  at foo\n  at bar"
  at bar              ┘
Exception: bang       ┐
  at baz              ├─ block 2
  at qux              ┘

A rule with pattern (?s)Exception.*at bar matches block 1 only, so matchCount is 1.

At the end of a tick, a partial block (fewer than N lines) is still evaluated rather than carried over to the next tick. That keeps alerting prompt, but it means block alignment restarts on every scan.

File-level checks

Independent of the regex rules, a log monitor can also watch the file’s behaviour:

Check Field Fires when
Staleness stalenessMinutes + stalenessSeverity The file has not changed for N minutes (a silent app)
Throughput burst throughputBurstKB + throughputBurstSeverity More than N KB written in one interval (a log storm)
Throughput silence throughputSilenceKB + throughputSilenceSeverity Fewer than N KB written in one interval (a stalled writer)

Set any of these to 0 to disable it.

Read volume and the per-tick line cap

New lines are streamed through the rules one at a time rather than buffered, so memory stays flat regardless of how large the file is or how many rules it carries. A single read feeds every rule on that file, so adding rules costs matching time, not I/O.

Each scan also caps how many new lines it reads, controlled by the agent-wide logMaxLinesPerTick setting (default 100000). The cap applies per file, per tick — with a glob path, each expanded file gets its own budget.

logMaxLinesPerTick is a safety valve, not a rate target. It does not mean “read 100,000 lines every minute”, and the agent does not attempt to spend the whole interval reading. On each tick it reads whatever has been appended since the last offset; if that backlog exceeds the cap, the first N lines are evaluated, the offset jumps to end-of-file, the rest of the backlog is skipped, and a warning is written to the agent log:

WARN log monitor: [/var/log/app.log] line cap 100000 reached between ticks, advancing offset to EOF

Skipped lines are never re-read on the next tick — they are gone from the rule counts. So the cap exists to stop a runaway writer from stalling the agent, and hitting it is a signal that either the scan interval is too long for that file’s rate, or the cap is set too low.

How fast is a scan, really?

Scan time is dominated by regular-expression matching, not by disk or by line count — reading and splitting lines is nearly free next to the cost of running each pattern over each line. So the answer is hardware-dependent, but far more pattern-dependent.

Reference figures below were measured on an AMD Ryzen 9 8940HX (Windows 11, Go 1.25) against a 100,000-line / 13.3 MB application log in the OS page cache. Treat them as orders of magnitude, not guarantees.

Pattern Example Time Throughput
No rules — read + line split only — 5 ms ~21 M lines/s
Literal substring ERROR 7 ms ~13 M lines/s
Char class + quantifier latency=[0-9]{3,}ms 13 ms ~8 M lines/s
Anchored literal ^[0-9T:.\-]+Z ERROR 23 ms ~4.4 M lines/s
Unanchored .* wildcard .*tenant [0-9]+.* 98 ms ~1 M lines/s
Word boundary \bERROR\b 149 ms ~670 k lines/s
Case-insensitive alternation (?i)fatal|panic|out of memory 428 ms ~230 k lines/s

Reading and splitting the file accounts for roughly 5 ms of the total in every row — everything beyond that is regex time. A careless pattern is ~60× more expensive than a careful one over the same data.

Monitor Time per tick Throughput Heap churn
1 rule (\bERROR\b) ~150 ms ~670 k lines/s ~14 MB
3 mixed rules ~0.6 s ~160 k lines/s ~14 MB
8 mixed rules ~1.6 s ~60 k lines/s ~15 MB
Any monitor, no new lines since last tick ~0.05 ms — ~78 KB

“Mixed” means a realistic blend including one \b boundary rule and one (?i) alternation rule. Cost scales close to linearly with the number of rules, weighted by each rule’s pattern cost.

Heap churn is transient garbage (one short-lived string per line), not resident memory — resident usage stays proportional to the number of rules, not to the file size. captureLines and joinLines add no measurable overhead.

The practical readings from those numbers:

  • An idle tick is free. A monitor whose file has not grown costs about 50 microseconds. Scanning many quiet files on a 1MIN schedule is not a concern.
  • The default cap is comfortable on a 1-minute schedule. Even a deliberately expensive 8-rule monitor chews through the full 100,000-line budget in under two seconds — a few percent of a 60-second interval, on one core. Budget roughly 0.5–2 seconds per 100,000 lines per file for a typical monitor.
  • Rule count and pattern quality are the levers. Prefer a literal or anchored prefix; avoid (?i) over an alternation when you can write the case variants explicitly, and avoid leading .*. Use excludePattern to drop noise, not extra rules.
  • Watch for the cap warning. If it appears, either shorten the schedule for that file (a 1MIN scan of a 100,000-line budget tolerates ~1,600 lines/second sustained), raise logMaxLinesPerTick, or split the glob so each file has its own budget.

Directory (folder) monitoring

A directory monitor scans one folder and evaluates rules on its contents. Use it to catch a spool that stops draining, a directory that fills up, or a drop-folder that never receives files.

Defining a directory monitor

Field Description
path Directory to scan
schedule Scan interval (1MIN … 60MIN)
recursive on walks sub-directories; off counts the top level only
alertOnMissing on alerts when the directory is absent, at alertOnMissingSeverity
alertOnMissingSeverity Severity for a missing directory (minor/major/critical, default critical)
rules List of content rules (below)

Each rule measures one property of the folder:

Field Description
name Rule label
metric fileCount (number of files) or totalSizeMB (total size in MB)
operator >, >=, <, <=, or =
threshold Value the operator compares against
severity minor, major, or critical
enabled off disables the rule

Test a monitor before you save

On the Sentinel Agent Settings page, each log-file and folder monitor has a ⚡ Test button. It runs that monitor’s rules against the real file or folder on the agent host right now and shows the outcome inline — per-rule match counts and status, captured sample lines, folder file-count / size, and staleness — so you can confirm a pattern or threshold behaves as intended without waiting for the next scheduled scan. The test reads from the start of the file (capped for responsiveness) and persists nothing. Throughput burst/silence checks are interval-based and are not evaluated by a one-shot test.

How a status is produced

For every kind, the flow is the same on each tick:

  1. The agent collects the value (a metric sample, a match count, a file size, …).
  2. It applies the rule’s operator and threshold.
  3. If breached, the result carries the rule’s severity (MINOR, MAJOR, or CRITICAL); otherwise OK.
  4. Results are stored for the dashboards and, when non-OK, drive alerts through the agent’s notification channels (email, Slack, Teams, PagerDuty) subject to the alert-timing settings.

See also

Translations