Skip to content
Studio DocsContact us

Operations

Monitoring, logs and alerts

See how your hosts and servers are doing, search their logs, scrape metrics into your own tools, and get told when something needs attention.

Metrics

  • Hosts report CPU (overall and per core), memory and disk every 30 seconds. The fleet view shows totals and picks out the hosts that need attention.
  • Servers report CPU, memory and player count. A reading that is not available is shown as "not reported", never as zero, so an empty server and an unreachable one never look the same.

Prometheus

Point your Prometheus at your organization's metrics endpoint with an API key that holds the monitoring:read scope. Scrape no more often than every 30 seconds; a faster scrape is refused with a Retry-After header.

The endpoint currently emits these server metrics, labelled with the group, server, host and game:

  • gamegrid_instance_up
  • gamegrid_instance_cpu_ratio
  • gamegrid_instance_memory_bytes
  • gamegrid_instance_players

Export

Server metrics can be exported as CSV or JSON for any range of up to 7 days. Exports are streamed, so a large range does not time out.

Logs

Server log files are shipped from the host when a log file rotates or the server exits, so they are not real time. For live output, use the server's console.

  • Search across your servers' logs over a window of up to 7 days, 50 results a page by default and up to 200.
  • Download logs in bulk as a bundle.

Alerts

Alerts appear in the studio app, where you can acknowledge and resolve them. Repeats of the same alert for the same resource are grouped rather than raised again.

  • Threshold rules. Create rules on host metrics: cpu_pct, cpu_total_pct, mem_used_mb and disk_used_gb.
  • Crash detection. Each host checks its servers every 10 seconds and raises an alert when one exits unexpectedly. Automatic restart after a crash can be switched on per server; it is off by default.

Alert types

Alerts marked "Yes" in the last column are also sent to GameGrid, because the fix is often ours.

AlertSeverityGameGrid also notified
instance.crashedwarningNo
instance.recovery_failedcriticalYes
instance.healthyinfoNo
host.offlinecriticalYes
host.onlineinfoNo
host.high_cpuwarningNo
host.high_memorywarningNo
host.disk_fullwarningNo
host.snapshot_unavailablewarningYes
hardware.release_queuedinfoNo
hardware.release_failedwarningYes
hardware.renewal_failedwarningNo
backup.failedwarningNo
backup.verification_failedcriticalYes
backup.quiesce_stuckcriticalYes
config.driftedwarningNo
version.disagreementwarningNo
schedule.revert_failedcriticalNo
schedule.reconcile_failedwarningNo
schedule.silence_release_failedwarningNo
schedule.run_failedwarningNo
steam.needs_guard_codewarningNo
steam.rate_limitedwarningNo
steam.sentry_invalidatedwarningNo
credit.lowwarning or criticalYes
credit.exhaustedcriticalYes
legal.subject_request_duewarningYes
legal.subject_request_overduecriticalYes
malware_detectedcriticalYes
scanner.unavailablewarningYes
instance.unresponsivewarningNo
rollup.failedwarningYes
partition.missingwarningYes
logs.shipping_degradedwarningNo
storage.quota_exceededwarningNo
bandwidth.quota_exceededwarningNo
webhook.failingwarningNo
notify.rate_limitedwarningYes
notify.delivery_failingcriticalYes
mail.complaintwarningYes

Where alerts go

Send alerts to contact points of these kinds: Email, Webhook, Discord, PagerDuty.

Routes decide which alerts reach which contact points. Quiet hours, silences (for example during maintenance) and digests keep the noise down.

For machine-readable events, use webhooks.

Talk to our studio team

Pricing and commercial terms: contact us for more information.

Hardware availability, contract terms and service commitments are agreed with each studio.