System, process, log & folder monitoring
The Sentinel Agent (Discovery Agent) watches a host from four angles: the whole system, individual processes, log files, and directories. Each is defined as a set of rules with severity thresholds, evaluated on a schedule, and turned into a Minor / Major / Critical status.
All four monitor kinds are stored as JSON in the agent and run as independent cron jobs. You define them two ways:
- Sentinel Agent Settings page — edit the monitors interactively in the agent’s Web UI.
- Component Config Builder — the guided wizard’s Custom Sentinel Agent flow walks you through log monitors, folder monitors, and the rest, then writes a ready-to-run
discovery_agent.json.
When a rule breaches its threshold the agent raises an alert at the rule’s severity and routes it through the configured notification channels.
The four monitor kinds
| Kind | Watches | Defined by |
|---|---|---|
| Global system | Host-wide metrics: CPU, memory, load, swap, disk, network, sockets | Per-metric Minor/Major/Critical thresholds |
| Process | Named/discovered processes: CPU, memory, threads, FDs, traffic | Per-process thresholds |
| Log file | One file or glob of files: pattern matches, staleness, throughput | A path + a list of rules |
| Directory | One folder: file count, total size | A path + a list of rules |
Schedules use the labels 1MIN, 5MIN, 10MIN, 15MIN, 30MIN, 60MIN. Each monitor runs as an independent job; if a scan takes longer than its interval, the overlapping tick is skipped rather than run concurrently, so a slow disk or a large directory can never stack up overlapping scans.
Global system monitoring
The agent samples host-wide metrics and compares each to its thresholds. Covered metrics include CPU, memory, load average (1/5/15 min), swap (in/out rate and usage), filesystem usage and inodes, disk I/O (read/write rate, queue, utilisation), network traffic in/out, TCP sockets, CPU I/O wait, and database connections.
Each metric has Minor / Major / Critical thresholds, set either as fixed values or computed automatically from a historical baseline (mean + standard deviation). Network traffic thresholds additionally carry an operator (> or <).
For the full field reference, see System Metric Thresholds and Threshold Configuration.
Process monitoring
Processes are monitored either by manual declaration (a process list) or by port scan (the process bound to a port). For each monitored process the agent tracks CPU, memory, thread count, file-descriptor count, and network traffic, each with its own severity thresholds.
For the full field reference, see Per-Process Metric Thresholds.
Traffic needs packet capture
0 and their thresholds can never fire — the Web UI marks those two pills and their charts as not measured rather than idle. Every other metric on this page comes from OS counters and is unaffected. See Bandwidth is always zero.
Log file monitoring
A log monitor watches one file — or a glob pattern such as /var/log/app-*.log — and evaluates a list of rules against the new lines appended since the last tick. The agent tracks a byte offset per file, so each scan reads only what was added; glob patterns pick up newly created files automatically.
Log rotation is handled automatically. Whether the file is rotated by rename-and-recreate (a new file takes the old name) or by copy-and-truncate (the file is emptied in place), the agent detects it and resumes from the start of the new file, so lines are neither silently skipped nor double-counted. Byte offsets are persisted, so a scan also resumes cleanly across an agent restart.
Defining a log monitor
| Field | Description |
|---|---|
path |
File path, or a glob pattern (*, ?, […]) for multiple files |
schedule |
Scan interval (1MIN … 60MIN) |
alertOnMissing |
on (default) alerts when the file is absent; off stays silent |
rules |
List of pattern rules (below) |
Each rule matches lines and alerts on how many matched:
| Field | Description |
|---|---|
name |
Rule label (shown in the alert) |
pattern |
Regular expression to match on each line |
excludePattern |
Lines also matching this are skipped |
operator |
>, >=, <, <=, =, or first |
threshold |
Match count the operator compares against |
severity |
minor, major, or critical |
captureLines |
Attach the first N matching lines to the alert text as samples (0 = none) |
joinLines |
Group lines into blocks of N and match the pattern against the whole block (0/1 = line by line) |
enabled |
off disables the rule without deleting it |
The first operator fires on the first occurrence regardless of count — useful for “alert the moment this ever appears” patterns like FATAL or panic.
captureLines — what lands in the alert
captureLines changes only what the alert shows, never what it counts. The rule always counts every match; capture keeps the first N of them, in file order, and appends them to the alert’s root cause after a | samples: marker, separated by ;:
log /var/log/app.log rule 'errors': 47 occurrence(s) of 'ERROR' > 10 | samples: ERROR db timeout on tenant 12 ; ERROR db timeout on tenant 44
Points worth knowing:
- First N, not last N. Matches 1…N are kept and the rest are discarded as the scan proceeds, so you get the oldest matches in the tick, not the most recent.
- Samples only appear when the rule fires. A rule that stays
OKproduces no root cause, so nothing is shown. - One line per sample. When
joinLinesis also set, only the first line of each matching block is captured — the sample is a pointer to the entry, not the entry itself. - Keep N small (2–5). The captured lines go into the alert body verbatim, and log lines can be long.
joinLines — how multi-line entries are matched
joinLines is not a “lines of context before the match” setting. It changes the unit that the regex is applied to: instead of testing one line at a time, the agent fills a buffer with N consecutive lines, joins them with newlines, and tests the pattern against that whole block. A block that matches counts as one match.
With joinLines = 3, a stack trace like this is evaluated as two blocks:
Exception: boom ┐
at foo ├─ block 1 → tested as "Exception: boom\n at foo\n at bar"
at bar ┘
Exception: bang ┐
at baz ├─ block 2
at qux ┘
A rule with pattern (?s)Exception.*at bar matches block 1 only, so matchCount is 1.
Three things to get right when using joinLines:
- Use the
(?s)flag. By default.does not match a newline, so a pattern meant to span the joined lines —(?s)Exception.*at bar— needs(?s)to cross the block’s internal newlines. - Counts are blocks, not lines.
thresholdis compared against the number of matching blocks. WithjoinLines = 5, a threshold of10means 10 blocks, i.e. up to 50 log lines. - Blocks do not slide. Grouping is fixed and non-overlapping (lines 1–3, 4–6, 7–9…), starting fresh at the beginning of each scan. An entry that straddles a block boundary can be split across two blocks and missed. Set N to the height of the entries you care about, and prefer a pattern that matches the entry’s first line plus a marker close to it.
At the end of a tick, a partial block (fewer than N lines) is still evaluated rather than carried over to the next tick. That keeps alerting prompt, but it means block alignment restarts on every scan.
File-level checks
Independent of the regex rules, a log monitor can also watch the file’s behaviour:
| Check | Field | Fires when |
|---|---|---|
| Staleness | stalenessMinutes + stalenessSeverity |
The file has not changed for N minutes (a silent app) |
| Throughput burst | throughputBurstKB + throughputBurstSeverity |
More than N KB written in one interval (a log storm) |
| Throughput silence | throughputSilenceKB + throughputSilenceSeverity |
Fewer than N KB written in one interval (a stalled writer) |
Set any of these to 0 to disable it.
Read volume and the per-tick line cap
New lines are streamed through the rules one at a time rather than buffered, so memory stays flat regardless of how large the file is or how many rules it carries. A single read feeds every rule on that file, so adding rules costs matching time, not I/O.
Each scan also caps how many new lines it reads, controlled by the agent-wide logMaxLinesPerTick setting (default 100000). The cap applies per file, per tick — with a glob path, each expanded file gets its own budget.
logMaxLinesPerTick is a safety valve, not a rate target. It does not mean “read 100,000 lines every minute”, and the agent does not attempt to spend the whole interval reading. On each tick it reads whatever has been appended since the last offset; if that backlog exceeds the cap, the first N lines are evaluated, the offset jumps to end-of-file, the rest of the backlog is skipped, and a warning is written to the agent log:
WARN log monitor: [/var/log/app.log] line cap 100000 reached between ticks, advancing offset to EOF
Skipped lines are never re-read on the next tick — they are gone from the rule counts. So the cap exists to stop a runaway writer from stalling the agent, and hitting it is a signal that either the scan interval is too long for that file’s rate, or the cap is set too low.
How fast is a scan, really?
Scan time is dominated by regular-expression matching, not by disk or by line count — reading and splitting lines is nearly free next to the cost of running each pattern over each line. So the answer is hardware-dependent, but far more pattern-dependent.
Reference figures below were measured on an AMD Ryzen 9 8940HX (Windows 11, Go 1.25) against a 100,000-line / 13.3 MB application log in the OS page cache. Treat them as orders of magnitude, not guarantees.
| Pattern | Example | Time | Throughput |
|---|---|---|---|
| No rules — read + line split only | — | 5 ms | ~21 M lines/s |
| Literal substring | ERROR |
7 ms | ~13 M lines/s |
| Char class + quantifier | latency=[0-9]{3,}ms |
13 ms | ~8 M lines/s |
| Anchored literal | ^[0-9T:.\-]+Z ERROR |
23 ms | ~4.4 M lines/s |
Unanchored .* wildcard |
.*tenant [0-9]+.* |
98 ms | ~1 M lines/s |
| Word boundary | \bERROR\b |
149 ms | ~670 k lines/s |
| Case-insensitive alternation | (?i)fatal|panic|out of memory |
428 ms | ~230 k lines/s |
Reading and splitting the file accounts for roughly 5 ms of the total in every row — everything beyond that is regex time. A careless pattern is ~60× more expensive than a careful one over the same data.
| Monitor | Time per tick | Throughput | Heap churn |
|---|---|---|---|
1 rule (\bERROR\b) |
~150 ms | ~670 k lines/s | ~14 MB |
| 3 mixed rules | ~0.6 s | ~160 k lines/s | ~14 MB |
| 8 mixed rules | ~1.6 s | ~60 k lines/s | ~15 MB |
| Any monitor, no new lines since last tick | ~0.05 ms | — | ~78 KB |
“Mixed” means a realistic blend including one \b boundary rule and one (?i) alternation rule. Cost scales close to linearly with the number of rules, weighted by each rule’s pattern cost.
Heap churn is transient garbage (one short-lived string per line), not resident memory — resident usage stays proportional to the number of rules, not to the file size. captureLines and joinLines add no measurable overhead.
The practical readings from those numbers:
- An idle tick is free. A monitor whose file has not grown costs about 50 microseconds. Scanning many quiet files on a
1MINschedule is not a concern. - The default cap is comfortable on a 1-minute schedule. Even a deliberately expensive 8-rule monitor chews through the full 100,000-line budget in under two seconds — a few percent of a 60-second interval, on one core. Budget roughly 0.5–2 seconds per 100,000 lines per file for a typical monitor.
- Rule count and pattern quality are the levers. Prefer a literal or anchored prefix; avoid
(?i)over an alternation when you can write the case variants explicitly, and avoid leading.*. UseexcludePatternto drop noise, not extra rules. - Watch for the cap warning. If it appears, either shorten the schedule for that file (a
1MINscan of a 100,000-line budget tolerates ~1,600 lines/second sustained), raiselogMaxLinesPerTick, or split the glob so each file has its own budget.
Directory (folder) monitoring
A directory monitor scans one folder and evaluates rules on its contents. Use it to catch a spool that stops draining, a directory that fills up, or a drop-folder that never receives files.
Defining a directory monitor
| Field | Description |
|---|---|
path |
Directory to scan |
schedule |
Scan interval (1MIN … 60MIN) |
recursive |
on walks sub-directories; off counts the top level only |
alertOnMissing |
on alerts when the directory is absent, at alertOnMissingSeverity |
alertOnMissingSeverity |
Severity for a missing directory (minor/major/critical, default critical) |
rules |
List of content rules (below) |
Each rule measures one property of the folder:
| Field | Description |
|---|---|
name |
Rule label |
metric |
fileCount (number of files) or totalSizeMB (total size in MB) |
operator |
>, >=, <, <=, or = |
threshold |
Value the operator compares against |
severity |
minor, major, or critical |
enabled |
off disables the rule |
Test a monitor before you save
How a status is produced
For every kind, the flow is the same on each tick:
- The agent collects the value (a metric sample, a match count, a file size, …).
- It applies the rule’s operator and threshold.
- If breached, the result carries the rule’s severity (
MINOR,MAJOR, orCRITICAL); otherwiseOK. - Results are stored for the dashboards and, when non-OK, drive alerts through the agent’s notification channels (email, Slack, Teams, PagerDuty) subject to the alert-timing settings.
See also
- Sentinel Agent component — discovery mechanisms, collected metrics, execution flow
- Sentinel Agent configuration — full settings and threshold reference
- Thresholds — how these metrics are evaluated into a status, and how to tune them
- Sentinel Agent HTTP API — programmatic access
- Monitor operations — probe-side monitors (URL, API, SNMP, WMI, …)