Triage Railway resource alerts before paging anyone

By General Input

When Railway fires a CPU, memory, or volume alert, an agent investigates the recent deploy and logs, classifies the cause, and posts a clear recommendation to Slack.

Integrations

  • Railway
  • Slack Bot
  • Linear

Type

Agentic Task

Categories

  • Engineering
  • Operations

Build me an agent workflow that triages Railway resource alerts before anyone gets paged.

Trigger: an incoming webhook from Railway. Railway sends a webhook when one of its resource alerts fires: CPU monitor alerts, RAM monitor alerts, and volume usage alerts (see https://docs.railway.com/guides/webhooks). The webhook payload identifies the project, environment, service, and the metric that crossed its threshold, plus the alert timestamp.

When an alert arrives, the agent should investigate before deciding what to do. Concretely:

1. Parse the webhook payload and pull out the project ID, environment ID, service ID, service name, metric type (CPU, RAM, or volume), threshold, current value, and alert timestamp.

2. Call Railway's List Deployments for that service and environment to see whether the spike correlates with a deploy in the last hour. Capture the deploy ID, status, and timestamp of the latest few deployments.

3. For the latest deployment, call Railway's Get Deployment Logs and Get HTTP Logs to gather runtime errors and recent request volume and status codes. Also call Get Environment Logs filtered to recent errors, so the agent has a wider view of what is happening across the environment around the alert time.

4. Classify the spike into exactly one of these buckets, grounded in concrete signals (deploy time vs alert time, error count delta, request volume delta, gradual rise over days), not vibes: deploy-induced regression, traffic spike, slow leak, or noisy threshold.

5. Recommend a concrete next step that matches the classification: rollback to a specific deploy ID, raise the threshold, scale replicas, or ignore. Pick one, do not list options.

Output 1: Post a single message to Slack using the Slack Bot Send a Message action. The message should include the service and environment, a severity tag (P1, P2, P3 based on production vs non-production and metric type), the classification, the recommended next step, and a short evidence block citing the deploy ID, error count, and request volume delta. Keep it scannable in one screen.

Output 2: If, and only if, the classification is deploy-induced regression or slow leak, also open a Linear issue using Linear Create Issue. The issue title should name the service and the cause, the description should include the log excerpts and the recommended action, and the assignee should come from a configurable service-to-owner mapping that the user sets up once.

Configurable inputs the workflow needs from the user: the Slack channel per environment (production vs non-production), the Linear team to file into, the service-to-owner mapping (service ID or name to Linear user email), and the time window the agent considers "recent" for deploys (default one hour) and for slow leaks (default seven days).

Important behavior: only one Slack message per alert (no duplicate retries), and skip filing a Linear issue if there is already an open Linear issue for the same service with the same classification in the last 24 hours. Keep the AI classification grounded in real signals from the logs, not guesses.

Related prompts

Explore more prompts
A brand asset library your marketing team actually searchesTurn Mailjet email clicks into ranked HubSpot follow-upsClean out the Looker dashboards and Looks nobody opensLiveKit live operations console for room moderationWake up dormant Keap leads with a researched reasonLiveChat coverage board for planning next week's shiftsPhone routing control panel for LiveKit voice agentsLinkedIn Ads budget pacing dashboard for every client accountGive your team Looker numbers without buying more seatsPause marketing emails to escalated customers, then restore them