Skip to main content
The Analytics page (/analytics) gives you a data-driven view of how your incident operations are performing over time. It is organized into five tabs — each tab loads its data independently so the page opens fast regardless of which tab you need. Use the date range picker at the top to change the reporting window. Preset options (7 days, 14 days, 30 days, 90 days) are available alongside a custom date range. Most tabs respect this window and your selected environment. A few cards are exceptions and stay unfiltered — the Cost tab’s budget-breach summary and episode usage card, the AI & Agents tab’s AI feedback tile and per-agent feedback table, and part of the Problems tab — all noted below. You can export the underlying data as CSV or JSON using the Export button in the top-right corner.

Overview tab

The Overview tab is the default landing surface. It focuses on business-outcome signals: are you resolving incidents quickly, and is the rate improving? KPI cards, in two tiers:
  • Total Incidents
  • MTTR (mean time to resolution)
  • AI Confidence
  • Auto Resolved (percentage resolved without human takeover)
  • Resolved (%)
  • Escalations
  • Open Problems
  • Problems Found
Incident trend chart — a day-by-day line chart showing incidents created vs. incidents resolved for the selected window. A growing gap between created and resolved is a signal that your team is falling behind. MTTR by severity — a bar chart breaking down average resolution time by severity level (P1–P4). P1 MTTR is the most operationally significant number here. Severity distribution — a breakdown of what proportion of incidents were P1, P2, P3, and P4 over the period.

Operations tab

The Operations tab surfaces the mechanics of how incidents moved through the lifecycle — where time was spent, what stopped them, and who was involved. Phase time chart — average time spent in each lifecycle stage (detection, triage, diagnosis, resolution, verification, closure). Use this to identify bottlenecks: if diagnosis is consistently the longest phase, that is a signal to improve runbook coverage or telemetry quality. Root causes — a breakdown of the most common root cause categories across resolved incidents in the period. Control mode chart — the proportion of incidents handled with the AI driving vs. a human in control. Control transfer — how often human intervention was triggered (including voluntary human takeovers), broken down by transfer reason. There’s no stage breakdown on this card — reason is the only dimension. Actor distribution — the split between AI, Human, and System actions across the period.

AI & Agents tab

The AI & Agents tab is the default landing surface for users in the Commander role. It focuses on how well the AI agents are performing — accuracy, recommendation outcomes, and approval latency. Two cards on this tab are exceptions to the date-range/environment filtering the rest of the page respects — both noted below. AI feedback (last 30d) — an aggregate thumbs-up / thumbs-down breakdown by surface (ECHO, SAGE, incident stages, RCA, and more) your team has submitted on agent outputs. This is a fixed trailing-30-day window with no binding to the date range picker or your selected environment. Per-agent feedback (last 30d) — a table breaking down the same thumbs-up / thumbs-down feedback per agent, with columns for Agent, Up, Down, Total, and Rate. Like the tile above, this is a fixed 30-day window regardless of the date range picker or selected environment. AI confidence distribution — a histogram of the agents’ root-cause confidence scores from the diagnosis stage specifically, not confidence at every recommendation point. A distribution skewed toward high confidence (80–100%) indicates the agents are working with good telemetry and runbook coverage. Recommendation outcomes — of all recommendations generated, how many are Approved (accepted, awaiting execution), Executed (applied to production), Rejected, Pending (still awaiting a decision), or Expired without action. Agent performance table — a per-agent breakdown (Incident Commander, Diagnosis Specialist, Resolution Specialist, Verification Specialist) showing analyses count, average confidence, recommendations count, and acceptance rate. Approval response time — how long it takes your team to respond to HITL approval requests, broken down by risk level (low / medium / high). Long approval delays on high-risk approvals are a signal to review your on-call escalation paths. Verification stats — total verifications, how many succeeded on the first pass, the resulting first-pass success rate, and how many required a re-diagnosis.

Problems tab

The Problems tab gives you a health check on your problem-management backlog and the quality of your alerting. Only part of this tab is unfiltered — the two cards below respect your selected environment and date range like the rest of the app: Problem intelligence — Conversion Rate (incident-to-problem), Total Downtime, Customers Affected, and SLA Breaches, plus an action-item completion ring and Duplicates / Related / Caused-By incident-link counts. Alert intelligence — average alerts per incident, total alerts received, and an Alerts by Source breakdown. The remaining cards on this tab are org-scoped and not filtered by environment or by the date range picker — they reflect the current backlog state, not a point-in-time window: Problem backlog health — SLA-based aging buckets (Overdue, Due soon, On track) for open problems, plus action-item completion and top recurring chains. Detection source mix — what proportion of incidents were detected automatically vs. reported by a user or a client vs. hybrid vs. unclassified. A higher automated share is the goal — user/client-reported incidents are a signal that monitoring coverage has a gap. Monitors proposed — how many new monitoring rules were proposed from problem action items, broken down by Staged, Approved, Pushed to provider, Rejected, and Push failed.

Cost tab

The Cost tab shows LLM spend for your organization. Its visibility is controlled purely by role — administrator or owner — regardless of billing mode.
Every production organization today is on bring-your-own-key (BYOK); demo/trial environments run on sureops-supplied capacity. A parallel managed AI billing mode is designed for the same non-dollar content described below but is not yet available for self-serve signup — see BYOK vs Managed LLM. What changes for an org on sureops-supplied capacity is the tab’s content, not its visibility: instead of the $ spend breakdown below, it sees an episode-usage card (or a “Cost analytics managed by sureops” card) in its place — the tab itself stays visible and clickable for any admin/owner.
Episode usage — two meters, Diagnosis and PR pre-mortems (the latter shown only when your plan includes pre-mortems), tracking episode-credit consumption against your plan’s included allowance. This card renders at the top of the tab, before the cards below, and has no period filter at all — the date range picker has zero effect on it, unlike every other card on this tab. For organizations on sureops-supplied capacity (demo/trial today; managed once launched), it appears in place of the dollar-spend cards below since there’s no $ spend to show. For BYOK organizations it still appears at the top of the tab, in addition to the dollar-spend analytics below it — every org is entitled to see its own episode-credit usage regardless of billing mode. Spend KPIs — four cards: total spend for the period, average cost per incident, LLM call count (with average cost per call shown as a hint line underneath), and total token usage. Budget breaches — reports how many times your configured budget was breached in the window, split into warn-level and enforce-level breaches, alongside your current enforcement posture (enforced, warn-only, or off). This isn’t a simple within/approaching/over-budget indicator — it’s a breach count plus policy state. Unlike the other Cost tab cards, this one is org-wide and doesn’t filter by your selected environment. Spend trend — a daily line chart of LLM spend over the period. Spend by stage — a bar chart showing which lifecycle stages drive the most LLM cost. Diagnosis is typically the most expensive stage. Spend by model — a table breaking down calls, tokens, and cost by AI model used. Spend by environment — a per-environment breakdown of incident count, call count, and cost, useful if your org runs multiple environments under one budget. Top expensive incidents — the 10 incidents that consumed the most LLM budget in the period, with links to their Incident Hubs.