Log in

Ops & Chief of Staff

Incident context assembled before the engineer opens their laptop

Alert details from PagerDuty, correlated metrics from Datadog, and service health from New Relic assembled into an incident brief with the likely cause, blast radius, and suggested rollback steps before the on-call engineer starts investigating.

Integrates with:

What changes

Before
With ZeroTwo
Context-gathering time
20-30 minutes opening tabs and cross-referencing timestamps
Incident brief ready within 2 minutes of the PagerDuty alert
Deploy correlation
Engineer manually checks deploy logs for recent changes
Recent deploys correlated with metric changes automatically
Blast radius visibility
Downstream services discovered as they start failing
Affected downstream services identified in the initial brief
Runbook access
Engineer searches Confluence for the right runbook mid-incident
Relevant runbook linked in the brief based on the service and failure type

Context-gathering time

Before

20-30 minutes opening tabs and cross-referencing timestamps

With ZeroTwo

Incident brief ready within 2 minutes of the PagerDuty alert

Deploy correlation

Before

Engineer manually checks deploy logs for recent changes

With ZeroTwo

Recent deploys correlated with metric changes automatically

Blast radius visibility

Before

Downstream services discovered as they start failing

With ZeroTwo

Affected downstream services identified in the initial brief

Runbook access

Before

Engineer searches Confluence for the right runbook mid-incident

With ZeroTwo

Relevant runbook linked in the brief based on the service and failure type

The first 30 minutes of every incident are wasted on context-gathering

PagerDuty fires at 2 AM. The on-call engineer opens Datadog, then New Relic, then checks the deploy log. They cross-reference timestamps, look for correlated service errors, and try to figure out what changed. Meanwhile, the service is down and the incident channel is filling up with 'any update?' messages.

The last P1 took 47 minutes to resolve. 28 of those minutes were spent gathering context: which service, what changed, when it started. The actual fix took 19 minutes.

How ZeroTwo builds the incident brief

1

Reads the P1 alert with service, escalation, and history

Slack

P1 alert fired at 2:14 AM. Service: payment-api. Escalation policy: on-call SRE. Alert details: 'payment-api response time exceeded 5s threshold for 3 consecutive checks.' Previous incident on this service: 11 days ago (bad config deploy).

2

Pulls metrics, error rates, and recent deploys for the affected service

Datadog

payment-api CPU spiked to 94% at 2:11 AM (baseline: 35%). Error rate jumped from 0.2% to 12.4%. Last deploy to payment-api: 2:08 AM by deploy-bot (commit: 'add retry logic to payment processor'). 2 other services showing elevated error rates downstream.

3

Checks service health, transaction traces, and upstream dependencies

New Relic

payment-api Apdex dropped to 0.12 (normal: 0.94). Transaction trace shows the retry loop in processPayment() averaging 8.3s per call (normal: 120ms). Upstream dependency map confirms downstream services are healthy independently but failing on calls to payment-api.

4

Assembles the incident brief with cause, blast radius, and next steps

Google

Incident brief: the payment-api deploy at 2:08 AM introduced a retry loop consuming CPU and cascading errors to downstream services. Suggested action: roll back the commit. Deploy diff, dashboards, and the rollback runbook linked.

Triggered on every P1 PagerDuty alert · Brief available within 2 minutes

Get started in under 10 minutes

1

Connect your tools

One-click OAuth for each integration. No API keys, no engineering.

2

Describe what you need

When a P1 alert fires in PagerDuty, pull the alert details, grab CPU, error rate, and recent deploys from Datadog, check service health in New Relic, and assemble an incident brief with the likely cause, affected services, and suggested rollback steps. Link the deploy diff and the relevant runbook.

3

It runs on schedule

Triggered on every P1 PagerDuty alert. Brief available within 2 minutes.

Frequently asked questions

Typically within 2 minutes. The bottleneck is API response times from Datadog and New Relic, not processing. Most briefs are ready before the on-call engineer finishes reading the PagerDuty notification.

ZeroTwo still provides the brief with metric anomalies, error rate changes, and service dependencies. If no recent deploy correlates, it says so and surfaces other potential causes like traffic spikes or upstream failures.

No. ZeroTwo assembles context and suggests next steps, but it does not roll back deploys or restart services. The on-call engineer makes every remediation decision.

PagerDuty for alert details and escalation history, Datadog for metrics and deploy events, and New Relic for service health and transaction traces. You can add additional sources like AWS CloudWatch or custom deploy tracking systems.

Yes. Each P1 alert generates its own brief independently. If two services fail at the same time, ZeroTwo also cross-references them to flag whether they share a common cause.

Related workflows

Stop doing the work your tools should do for you.

Set it up once. ZeroTwo runs it every time.