Operations
Monitoring, logs and alerts
See how your hosts and servers are doing, search their logs, scrape metrics into your own tools, and get told when something needs attention.
Metrics
- Hosts report CPU (overall and per core), memory and disk every 30 seconds. The fleet view shows totals and picks out the hosts that need attention.
- Servers report CPU, memory and player count. A reading that is not available is shown as "not reported", never as zero, so an empty server and an unreachable one never look the same.
Prometheus
Point your Prometheus at your organization's metrics endpoint with an API key that holds the monitoring:read scope. Scrape no more often than every 30 seconds; a faster scrape is refused with a Retry-After header.
The endpoint currently emits these server metrics, labelled with the group, server, host and game:
gamegrid_instance_upgamegrid_instance_cpu_ratiogamegrid_instance_memory_bytesgamegrid_instance_players
Export
Server metrics can be exported as CSV or JSON for any range of up to 7 days. Exports are streamed, so a large range does not time out.
Logs
Server log files are shipped from the host when a log file rotates or the server exits, so they are not real time. For live output, use the server's console.
- Search across your servers' logs over a window of up to 7 days, 50 results a page by default and up to 200.
- Download logs in bulk as a bundle.
Alerts
Alerts appear in the studio app, where you can acknowledge and resolve them. Repeats of the same alert for the same resource are grouped rather than raised again.
- Threshold rules. Create rules on host metrics:
cpu_pct,cpu_total_pct,mem_used_mbanddisk_used_gb. - Crash detection. Each host checks its servers every 10 seconds and raises an alert when one exits unexpectedly. Automatic restart after a crash can be switched on per server; it is off by default.
Alert types
Alerts marked "Yes" in the last column are also sent to GameGrid, because the fix is often ours.
| Alert | Severity | GameGrid also notified |
|---|---|---|
instance.crashed | warning | No |
instance.recovery_failed | critical | Yes |
instance.healthy | info | No |
host.offline | critical | Yes |
host.online | info | No |
host.high_cpu | warning | No |
host.high_memory | warning | No |
host.disk_full | warning | No |
host.snapshot_unavailable | warning | Yes |
hardware.release_queued | info | No |
hardware.release_failed | warning | Yes |
hardware.renewal_failed | warning | No |
backup.failed | warning | No |
backup.verification_failed | critical | Yes |
backup.quiesce_stuck | critical | Yes |
config.drifted | warning | No |
version.disagreement | warning | No |
schedule.revert_failed | critical | No |
schedule.reconcile_failed | warning | No |
schedule.silence_release_failed | warning | No |
schedule.run_failed | warning | No |
steam.needs_guard_code | warning | No |
steam.rate_limited | warning | No |
steam.sentry_invalidated | warning | No |
credit.low | warning or critical | Yes |
credit.exhausted | critical | Yes |
legal.subject_request_due | warning | Yes |
legal.subject_request_overdue | critical | Yes |
malware_detected | critical | Yes |
scanner.unavailable | warning | Yes |
instance.unresponsive | warning | No |
rollup.failed | warning | Yes |
partition.missing | warning | Yes |
logs.shipping_degraded | warning | No |
storage.quota_exceeded | warning | No |
bandwidth.quota_exceeded | warning | No |
webhook.failing | warning | No |
notify.rate_limited | warning | Yes |
notify.delivery_failing | critical | Yes |
mail.complaint | warning | Yes |
Where alerts go
Send alerts to contact points of these kinds: Email, Webhook, Discord, PagerDuty.
Routes decide which alerts reach which contact points. Quiet hours, silences (for example during maintenance) and digests keep the noise down.
For machine-readable events, use webhooks.