Monitor server status

Monitor server resource usage and performance in real-time.

Monitoring dashboard

How to access

  1. Open Metrics from the workspace app
  2. Click the server you want to look at
  3. View real-time metric charts

Monitoring metrics

Server charts (CPU, memory, disk, disk I/O, and network) and metric alert rules are part of the Metrics extension. Free plans can’t enable it. Essentials and Enterprise workspaces can turn it on from Workspace settings > Integrations > Extensions. A workspace that was already on a paid plan has it on; a newly created workspace starts with it off.

With the extension on, the dashboard charts:

  • CPU usage: Overall CPU utilization
  • Memory usage: System memory utilization
  • Disk usage: Disk usage per partition
  • Disk I/O: Read/write speeds per disk (Peak/AVG)
  • Network traffic: Network usage per interface (bps/pps)

New data arrives about once a minute while the extension is on.

With the extension off, the charts and metric alerts say the extension isn’t enabled—or that your plan can’t enable it—instead of showing empty charts.

Every plan keeps two things regardless of the extension: server online/offline status and the offline alert, and the disk usage behind the core disk-usage alert—normally just the server’s own volume, every volume if the agent doesn’t identify one. The latest reading is on the server’s Overview tab under Agent information, as Volume usage, on every plan and whether or not the extension is on. The same value is in the server’s collection profile for API callers.

From 2.37.0, the metrics overview reads every server’s latest values through the latest metrics endpoint, which any script or integration can call the same way.

Network interfaces

From alpamon v2.7.3, upgrading the agent doesn’t remove any interface from a server’s interface list. What changes is the network traffic chart: container-style virtual interfaces—one half of a veth pair, a macvlan or ipvlan interface, a tun or tap device, and dummy interfaces—stop reporting traffic, so the chart no longer fills up with short-lived container interfaces. They stay in the interface list; only their chart goes empty. If that would leave a server with no interface reporting traffic at all, such as an agent running inside a container, the agent reports traffic for them anyway.

Loopback and interfaces with no hardware address, such as WireGuard, OpenVPN tun, or PPP links, were never shown and can’t be added with the setting below.

To see traffic for a specific virtual interface again, add it to include_virtual under [interface] in the agent’s configuration file, then restart the agent:

[interface]
include_virtual = veth0, cali*

To also drop those interfaces from the interface list, not just their chart, set exclude_virtual_from_inventory = true under [interface]—but only against an Alpacon server on 2.37.2 or later. On an older server, removing an interface this way deletes its traffic history along with it.

On macOS and Windows, the agent still decides which interfaces report traffic by name, as it always has.

Viewing charts

Hover over the chart line to view detailed information in a tooltip.

Select time range

Available time ranges vary by plan. See Plans & billing for what each plan unlocks.

Select time range from chart top:

  • Last 1 hour
  • Last 24 hours (1 day)
  • Last 7 days (1 week)
  • Last 1 month
  • Last 3 months
  • Last 1 year
  • Custom range

A range is offered only when it fits inside what your plan keeps, so Essentials stops at the last month and Enterprise offers all of them.

Refresh charts

You can set the refresh interval in minutes in the Refresh field.
Click the arrow icon button on the right to refresh charts immediately.

Metric alerts

Metric alerts are also part of the Metrics extension—turn it on the same way as the charts above. A rule watches one metric and raises an alert when it crosses a threshold you set.

How rules fire

Alpacon checks every rule about once a minute, using the data the agent has already sent. A rule fires when the value has stayed past the threshold for the duration you set, or on the first sample if the duration is zero. It clears only when the value comes back past the recovery threshold, or past the threshold itself if you didn’t set one—so a value hovering right at the line doesn’t flip the alert on and off.

A positive duration also has a floor: it has to span at least two collection intervals for the target you’re watching, since anything shorter can’t reliably hold a second sample and would behave like an instant rule while reading as a sustained one. 0 (fire on the first sample) isn’t affected—see what a rule can carry for where duration is set.

The default CPU usage and memory usage rules that come with every workspace need the value to stay at or above 80% for about 5 minutes before they fire, and they clear once it drops back below 75%. The default disk usage rule fires as soon as one sample reaches 80%, and clears below 78%.

You can alert on a value that’s too high or too low. For disk usage, disk I/O, and network rules, each disk or interface raises its own alert, and you can limit a rule to just one device instead of every one the server reports.

A rule can also alert when a server stops reporting a metric for longer than you set. Alerts stay open while a server is silent—they don’t clear just because the data stopped.

Each rule also carries a severity—critical, warning, or info—shown on the alert it raises, and severity decides whether that alert can email or post to Slack. See what a rule can carry for how each one behaves.

Attach an alert to a server

  1. Open Metrics from the workspace app
  2. Select the server and go to the Rules tab
  3. Click Assign rule
  4. Select from workspace-created alerts

Alerts have to be created in the workspace first, on the Metric alerts screen of the Metrics app.

Add rule

Created on the Metric alerts screen of the Metrics app.

  1. Click New rule
  2. Pick the target. For CPU usage, memory usage, and disk usage, the form opens with the values of the default rules every workspace starts with. For disk I/O and network targets, enter the threshold yourself
  3. Change what you need under Condition, Scope, Delivery, and Advanced
  4. Save

Every field is listed in what a rule can carry.

The list also says, next to its own description, how many servers have no performance rule attached—as a link straight to them—so long as there’s at least one.

Note: Only one default rule can be created per target

What are default rules? Rules automatically applied when new servers are registered. Every workspace starts with three: CPU usage and memory usage fire once the value stays at or above 80% for about 5 minutes and clear below 75%, and disk usage fires on the first sample at or above 80% and clears below 78%.

Edit rule

Click a rule’s name to open its own page, which has two tabs. Rule says what the rule does in one sentence, then repeats it field by field, and Update opens the same form New rule uses. Applied servers lists every server this rule is attached to—a server exempted from the rule by its own override still appears there. The Actions menu on the list row edits and deletes without leaving the list.

  • Cannot delete the only default rule for a target

Who can edit or delete a rule

A workspace admin can edit or delete any rule. So can anyone whose role carries rule-editing and rule-deleting rights—the built-in operator role has both, member does not. Creating a rule doesn’t give its creator any special access to it by itself.

An admin can also give a specific person, group, or service token the right to edit and delete just one rule, without changing their overall role—ask a workspace admin if your role doesn’t already cover a rule you need to manage. This grant can’t be made from the rule’s own page yet.

Some older rules still let the person who created them edit and delete that one rule, from a grant Alpacon made automatically at creation; Alpacon is cleaning those up over time.

The default CPU usage and memory usage rules can be edited by anyone with rule-editing rights, the same as any other rule. The disk usage rule that comes with every workspace is different: only a workspace admin can edit it, turn another rule into it (or it into an ordinary rule), add, change, or remove a per-server override on it, or detach it from a server. Attaching it back to a server isn’t restricted this way, and anyone who can already see the rule can still preview it.

No default rule, including the disk usage rule, can be deleted while it’s still marked as default.

What a rule can carry

The console form sets all of these, grouped as Condition, Scope, Delivery, and Advanced. The same fields are available through the alert rules API:

FieldWhat it doesValidation
NameIdentifies the ruleMust be unique in the workspace
TargetWhich metric the rule watches: CPU usage, memory usage, disk usage, disk I/O (peak or average read/write), or network traffic (peak or average input/output, bps or pps)—
ThresholdThe value that raises the alert—
DirectionAlert when the value is at or above the threshold, or at or below it—
DurationHow long the value has to stay past the threshold before the alert fires0 fires on the very first sample; a positive value must be at least two collection intervals for the target
Recovery thresholdThe value the metric has to come back to before the alert clearsMust be on the recovering side of the threshold: below it when the Direction is at or above, above it when the Direction is at or below. Leave it empty and the threshold itself clears the alert
No-data windowRaise a separate alert if this metric stops arriving for longer than thisMust be at least as long as how often this metric is collected, and at most 24 hours
DeviceLimit the rule to one disk or network interfaceOnly for disk usage, disk I/O, and network rules—CPU and memory apply to the whole server and can’t be scoped to a device
SeverityThe severity the raised alert carries, and which of the two destinations below can actually fire: critical can email and post to Slack, warning can only post to Slack, info reaches neitherDefaults to warning
Email destinationWho the alert email goes to, narrowed from everyone who can already see the server: everyone who can see it, workspace admins only, server group members and the owner, or no individual alert email. Only takes effect on a critical-severity ruleDefaults to everyone who can see the server
Slack channel postWhether a raise and its resolution post to the workspace’s connected Slack channelDefaults to on
Default ruleAutomatically attached to every server registered afterwardOnly one default rule per target and device

Severity decides which of the two destinations above can actually fire—the fields themselves only say who or what within that. A rule set to critical severity sends mail and posts to Slack exactly as those two fields say. A rule set to warning severity (the default) never emails, whatever the email destination says—it posts to Slack only if that field is on. A rule set to info severity does neither. Whatever the severity, every alert is still recorded in the console and your notification bell, reaches any webhook subscribed to the rule’s events, and is included in the hourly open alert digest sent to workspace admins while it’s open. Set a rule to critical if it needs to keep emailing.

A server can also depart from a rule’s threshold, recovery threshold, or duration on its own, or be exempted from the rule entirely—see Per-server overrides. The email and Slack destinations above, and severity, are workspace-wide only.

Per-server overrides

A server can depart from a workspace rule without changing it for everyone else: give that server its own threshold, recovery threshold, or duration, or exempt it from the rule entirely. The disk usage rule that comes with every workspace is the one exception—see Who can edit or delete a rule.

Set it on the server’s Rules tab: open the row’s actions and choose Set per-server override. The form opens holding the rule’s own values, so what you see is what the server does today. A box left at the rule’s value keeps that field following the rule, and a later change to the rule still reaches this server; type a different number and only that field departs. Clear Apply this rule to this server to exempt the server from the rule entirely, and Go back to the rule’s values removes the override.

The Applied column on that tab says which of the three each rule is in: Rule values, Overridden, or Not applied.

The same settings are available through the rule overrides API, for any server you can manage.

Receiving alerts

A monitoring alert—whether it’s a threshold breach or a metric going silent—shows up in the console’s alert list and on the dashboard, and reaches any subscribed webhook, the moment the rule fires. Email and the Slack post (if connected) wait a little longer—about two minutes—to confirm the condition is still true before interrupting anyone. If it clears before then, no email or Slack message goes out at all for that alert: the alert still shows as resolved in the console’s alert list, and a webhook subscribed to the resolved event still receives it.

Once an alert has stood long enough to be emailed or posted to Slack, where it goes depends on the rule’s severity. A critical alert reaches whoever the rule’s email destination names, and, if the rule’s Slack channel post is on, the workspace’s connected Slack channel; several alerts that fire together are combined into one summary. A warning alert—the default for a new rule—never emails, whatever its email destination says; it posts to Slack only if that setting is on, and either way it appears in the hourly open alert digest sent to workspace admins while it’s open. An info alert reaches neither channel. A critical rule emails everyone who can see the server by default, and can be narrowed to workspace admins only, the server’s group members and owner, or no individual alert email—see what a rule can carry for those two fields, and set a rule’s severity to critical if you need it to keep emailing. See How you’re notified for who counts as everyone who can see the server, why the email names you, and how to turn off your own email.

For a single alert, the email and Slack subject carry what the rule saw—for example, CPU usage 92% ≥ 80% for 5 minutes on web-01 for a rule with a duration set, or CPU usage stopped reporting on web-01 for a metric that’s gone silent. Several alerts fired together get a short summary subject instead.

When an alert you were emailed or Slack-notified about clears, the same people get a follow-up saying why: the metric recovered, the server started reporting data again, the rule changed or no longer covers the server, the server was exempted from the rule, or the server is no longer being managed.

A webhook, or any other notification channel, can subscribe directly to a rule crossing and clearing instead of reading email or Slack, with neither wait. See Metric threshold events for the two event types and their payload.

Alert history

Alert history in the Metrics app lists the metric alerts that have already cleared: what fired, its Rule, on which server, its severity and device, when it was raised and when it cleared, how long it stood, and its Cause—why it closed. The Rule column links to the rule’s own page—see Edit rule—and shows the alert’s condition, Threshold or Metric went silent, underneath the name; if the rule was deleted, or none was recorded, it reads No rule instead of a link. Cause uses the same values as the cause field on a resolved event: the metric recovered, the server started reporting again, the rule was deleted or the server was detached from it, the rule changed or no longer covers this server or device, a per-server override exempted the server, the server is no longer managed, no rule covers it anymore, a per-device alert superseded it, the device is no longer reported by the server, or the condition no longer applies for a cause that couldn’t be more specifically identified. An alert resolved before this column existed shows no cause. Alerts still standing are on the dashboard rather than here. The rule’s own page links back here too, with an Alert history link that filters this list to just that rule.

The Delivered column shows what the alert’s raise actually sent. A channel that delivered gets its own chip—Email · 3, or Slack with no count, since a channel post addresses a room rather than a list of people. A channel whose send failed shows Email · Failed to send (or Slack · Failed to send) instead, with a tooltip noting that the send was attempted and failed. A raise that sent nothing shows one muted chip instead, with a tooltip naming why: it’s still inside its hold, the condition cleared before the hold ended, the rule’s own destinations kept it off every channel, or it was dispatched but no delivery was recorded. An alert type that never sends an interrupt shows Not recorded.

History rows are kept for 30 days after the alert resolves; open alerts are never purged, however long they’ve stood. Metric alerts raised by an extension rule—anything beyond the seeded core disk-usage rule—are left out of this list, the dashboard, and the digest while the Metrics extension is off. The core disk-usage rule’s own alerts stay visible either way.

Click Filter to narrow the list to one Rule, or use Date range to narrow it to when the alert was raised—Today, Last 7 days, Last 30 days, or a custom range.

Offline alert

Alpacon raises an offline alert when a server has been offline for about 5 minutes. It clears on its own as soon as the server reconnects.

The alert stays quiet right after you restart or upgrade the agent—that’s expected downtime, not an outage. It still appears if the server hasn’t reconnected by the time below:

What you didAlert appears if the server isn’t back by
Restarted the agentAbout 15 minutes
Upgraded the agentAbout 40 minutes

A reboot, shutdown, or OS package upgrade you run yourself inside a work session doesn’t get this automatic suppression—mute the server’s alerts first if you’re planning one.

Stopping or terminating a cloud instance through your cloud provider doesn’t raise the alert either.

How you’re notified

When a server has been offline for about 5 minutes, everyone who can see that server’s alerts—workspace admins, members of a group the server belongs to, and the server’s owner—gets an email, and a message posts to the workspace’s Slack channel if one is connected. It still appears in the console’s alert list too.

If several servers go offline at once, you get one summary email and one Slack post listing all of them, instead of a separate message per server.

When the server reconnects, the same people get a follow-up email that says why the alert closed—the agent reconnected, or one of the other reasons below—and how long it was open. If the alert was posted to Slack, a reply in that thread confirms the alert cleared and how long the server was offline; it doesn’t repeat the reason, so check the email for that. If the alert closed another way instead—offline alerts were turned off for that server, the server was disabled or removed from the workspace, or your cloud provider shows the instance stopped—the follow-up email says that instead.

If a server keeps disconnecting and reconnecting—three or more times within an hour—Alpacon pauses the per-disconnection email and Slack post and sends one notice instead, subject web-01 has an unstable connection, saying how many times it’s dropped in the last hour. The alert list, dashboard, and bell keep recording every individual disconnect and reconnect; only the mail and Slack messages pause. Once the server has stayed connected for 30 minutes, one email titled Resolved: web-01 has a stable connection again goes out—with a Slack reply in the same thread if the notice was posted there—and per-disconnection mail resumes. If it never reconnects, the standing outage is emailed as an ordinary offline alert once 30 minutes have passed with the link still down. Turning off the offline alert for the server (below) skips this too.

Every alert email ends with a line saying why you received it—workspace admin, member of a group the server belongs to, or the server’s owner—followed by a link to Preferences > Notifications. Your own notification preferences apply to this alert too—muting email still leaves the Slack post in place for the channel.

Turn it off for a server

Turn off the offline alert for a server that’s meant to go offline on its own—an intermittently powered host, or one that shuts itself down after finishing a job, for example.

In the server’s edit form, turn off Alert when this server goes offline. If an alert is currently open for that server, it isn’t closed the moment you turn the toggle off; it closes on the next check, within about 5 minutes, and the follow-up email explains why.

You can also set it through the API:

curl -X PATCH "https://your-workspace.us1.alpacon.io/api/servers/servers/{server_id}/" \
  -H "Authorization: token=\"alpat-xxxxxxxxxxxxxxxxxx\"" \
  -H "Content-Type: application/json" \
  -d '{"offline_alert_enabled": false}'

See Servers API for the field reference.

Mute alerts

Turn off a server’s alert interrupts for a while—useful during planned work like a disk swap, so a server you already know is affected doesn’t keep paging you.

Mute a server

Open Mute alerts from a server’s detail page, or select one or more servers on the server list and choose Mute alerts from the bulk actions. Pick how long: 1 hour, 4 hours, or until a set time you choose—up to 7 days from now. You can also add a reason, up to 255 characters.

While a server is muted, a badge naming the end time—for example, Muted until 18:00—appears next to it everywhere it shows up: the server list, its card, and its detail page. The badge disappears on its own once the mute ends.

To turn a server’s alerts back on early, open the same control and choose Unmute alerts.

While a server is muted

Alerts are still raised and listed normally—in the bell and the alert list, and, once resolved, in Alert history for a metric alert or the server’s own offline alert history for the offline alert—but nothing interrupts anyone about them until the mute ends: no email, Slack message, push notification, or webhook. This covers every alert the server can raise, including metric threshold alerts and the hourly open alert digest. In Alert history, the Delivered column shows a metric alert held this way as Held by mute.

When the mute ends

If an alert’s condition is still open once the mute ends, its interrupt goes out then—held, not dropped. If the condition cleared while the server was muted, nothing is ever sent for it. A resolution notice for an alert that had already gone out before the mute began isn’t held either way—it’s delivered as usual.

Who can mute a server

Muting or unmuting a server takes the same permission as any other write to it: its owner, an owner or manager of an assigned group, or Staff/Superuser. A service token also needs the server:mute scope. The Alpamon agent can’t mute or unmute a server.

Every mute and unmute is recorded in the activity log, with who did it, when the mute ends, and the reason if one was given.

How this differs from a maintenance window

A maintenance window and a mute both quiet a server’s alerts for a while, but not the same way:

  • A maintenance window suppresses the alert itself, so nothing is raised and no row is written. A mute still raises and lists the alert—only its interrupts are held.
  • A maintenance window doesn’t affect metric threshold alerts, only the offline alert and a metric going silent. A mute holds metric threshold alerts too.
  • A maintenance window defers an alert’s offline resolution notice until the window ends. A mute doesn’t hold the resolution notice for an alert that was already announced before it began—that notice goes out right away.

A subscribed webhook behaves the same either way for the offline alert and a metric going silent: nothing while the window or mute is active, one event once it ends. A metric threshold alert’s webhook isn’t affected by a maintenance window at all—only a mute holds it.

View logs

Agent logs

  1. Server detail > Logs tab
  2. Check Alpamon agent logs in Recent logs section

Log content:

  • Agent start/stop
  • Connection status changes
  • Error and warning messages

Backhaul history

Check data transmission history in Recent backhauls section:

  • Transmission time
  • Transmission status
  • Data size

Server overview

View comprehensive server information in Overview tab.

System information:

  • Agent: Alpamon version and status
  • OS: Operating system and distribution
  • Hardware: CPU, memory, disk specs
  • Operating: Uptime, boot time

Additional information:

  • System packages: List of installed packages
  • Network interfaces: Active network interfaces

Activity history

View command execution history in Activity tab.

Command history:

  • Execution time
  • Executing user
  • Executed command
  • Execution result

Troubleshooting

Metrics not displaying:

  • Verify the Metrics extension is enabled (Workspace settings > Integrations > Extensions)
  • Verify server is in Connected status
  • Confirm agent is running normally
  • Check error messages in Logs tab
  • To see exactly what a server collects and why a metric might be missing, check its collection profile

Alerts not being sent:

  • Check whether the server is muted
  • Verify metric alerts are configured correctly
  • Check that you can see the server (workspace admin, a member of a group it belongs to, or its owner)
  • For a metric alert, check the rule’s severity—only a critical-severity rule sends individual email; a warning-severity rule posts to Slack (if that’s on) and appears in the hourly digest instead, and an info-severity rule sends neither
  • For a critical-severity rule, check the rule’s email destination and Slack channel post setting—it may be narrower than everyone who can see the server, or turned off
  • Check spam folder
Last updated: