# sureops ## Docs - [Budgets & cost](https://docs.sureops.ai/next/administration/budgets-and-cost.md): Track LLM spend per environment, set budget limits, and receive cost alerts before you overspend - [Deploy identity & recovery cooldown](https://docs.sureops.ai/next/administration/deploy-identity.md): Configure the service account sureops uses to execute remediation actions, and set minimum wait times between recovery attempts - [Notifications & channels](https://docs.sureops.ai/next/administration/notifications.md): Configure where sureops sends alerts, approval requests, and status updates — per event, per channel, per escalation path - [Roles & capabilities](https://docs.sureops.ai/next/administration/roles-and-capabilities.md): Five roles from Observer to Owner — what each can do, and how to pick the right one - [Team management](https://docs.sureops.ai/next/administration/team-management.md): Invite members, assign roles, and remove access from your sureops organization - [BYOK vs Managed LLM](https://docs.sureops.ai/next/ai-agents/byok-vs-managed.md): Choose between sureops-managed LLM keys and bringing your own provider API keys. - [Incident Pipeline Agents](https://docs.sureops.ai/next/ai-agents/incident-pipeline-agents.md): The four agents that run during an active incident — Commander, Diagnosis Specialist, Resolution Agent, and Verification Agent. - [Per-Agent LLM Config](https://docs.sureops.ai/next/ai-agents/per-agent-llm-config.md): Configure which LLM model each agent uses, with per-environment overrides. - [Post-Closure Agents](https://docs.sureops.ai/next/ai-agents/post-closure-agents.md): The agents that run after an incident closes — Problem Investigator, Fix-PR Agent, Pre-mortem Agent, and Curator. - [Skills](https://docs.sureops.ai/next/ai-agents/skills.md): Reusable tool bundles that extend what agents can do — and how to configure them. - [Control modes](https://docs.sureops.ai/next/getting-started/core-concepts-control-modes.md): Supervised and Self-Driving — how sureops operates alongside your team - [HITL approvals & confidence](https://docs.sureops.ai/next/getting-started/core-concepts-hitl.md): Human-in-the-loop gates, confidence surface, and decision trace - [The 6-stage incident lifecycle](https://docs.sureops.ai/next/getting-started/core-concepts-lifecycle.md): Detection → triage → diagnosis → resolution → verification → closure — what happens at each stage - [ECHO vs SAGE](https://docs.sureops.ai/next/getting-started/echo-vs-sage.md): Organization assistant vs per-incident advisor — when to use each - [Quickstart](https://docs.sureops.ai/next/getting-started/quickstart.md): Evaluate sureops with a demo environment in 5 minutes - [What is sureops](https://docs.sureops.ai/next/getting-started/what-is-sureops.md): Trust-first AI SRE — from Supervised to Self-Driving incident response - [ArgoCD](https://docs.sureops.ai/next/integrations/argocd.md): Connect ArgoCD for GitOps deployment context and rollback actions during resolution - [GitHub](https://docs.sureops.ai/next/integrations/github.md): Connect GitHub for source context, fix PRs, and change-tracking issues - [Grafana](https://docs.sureops.ai/next/integrations/grafana.md): Connect Grafana for dashboard context, metric queries, and unified alerts - [How integrations work](https://docs.sureops.ai/next/integrations/how-integrations-work.md): MCP-based integration model — vendor-hosted and customer-provided - [Kubernetes](https://docs.sureops.ai/next/integrations/kubernetes.md): Connect Kubernetes for pod inspection, cluster events, and deployment state during incidents - [Loki](https://docs.sureops.ai/next/integrations/loki.md): Connect Loki directly for LogQL log search during diagnosis - [Prometheus](https://docs.sureops.ai/next/integrations/prometheus.md): Connect Prometheus directly for PromQL metric queries during diagnosis and verification - [Slack](https://docs.sureops.ai/next/integrations/slack.md): Connect Slack for incident channels, team notifications, and HITL approvals - [Tempo](https://docs.sureops.ai/next/integrations/tempo.md): Connect Tempo for distributed trace search and TraceQL queries during diagnosis - [Webhooks & alert routing](https://docs.sureops.ai/next/integrations/webhooks-alert-routing.md): Ingest alerts from any source via webhook - [Adopt: connect your stack](https://docs.sureops.ai/next/onboarding/adopt-connect-your-stack.md): The 5-step guide to connecting your own infrastructure - [Demo vs Adopt environments](https://docs.sureops.ai/next/onboarding/demo-vs-adopt.md): Two paths to using sureops — sample environment and your own infrastructure - [Guided onboarding journey](https://docs.sureops.ai/next/onboarding/guided-onboarding-journey.md): The 13-step in-app onboarding tour - [Proactive detection](https://docs.sureops.ai/next/onboarding/proactive-detection.md): How sureops detects anomalies before alerts fire and how to tune it - [Confidence & Audit Trace](https://docs.sureops.ai/next/operating-modes/confidence-and-audit-trace.md): How sureops surfaces its reasoning and decision history so you can calibrate trust over time. - [HITL Gates](https://docs.sureops.ai/next/operating-modes/hitl-gates.md): When sureops pauses to ask for your approval, what those requests look like, and how to configure timeout behavior. - [Security & Compliance](https://docs.sureops.ai/next/operating-modes/security-and-compliance.md): Multi-tenant isolation, data flow, LLM data handling, and compliance posture. - [Supervised vs Self-Driving](https://docs.sureops.ai/next/operating-modes/supervised-vs-self-driving.md): How to configure whether sureops drives autonomously or pauses at every stage for your approval. - [Incident channels](https://docs.sureops.ai/next/problem-management/incident-channels.md): Per-incident Slack war rooms — automatic channel creation, status updates, and HITL approvals in Slack - [Recurrence chains](https://docs.sureops.ai/next/problem-management/recurrence-chains.md): How sureops links recurring incidents to the same root problem and lets you track re-opening frequency - [Tracker sync](https://docs.sureops.ai/next/problem-management/tracker-sync.md): Automatically push problem records to your connected issue tracker and keep them in sync - [Connect your IDE](https://docs.sureops.ai/next/reference/connect-your-ide.md): Use sureops documentation via MCP in Cursor, VS Code, and other AI-assisted IDEs - [Give feedback](https://docs.sureops.ai/next/reference/give-feedback.md): Report bugs, request features, ask questions, and improve the sureops docs - [Glossary](https://docs.sureops.ai/next/reference/glossary.md): Key terms and concepts in sureops, defined - [Privacy & AI use](https://docs.sureops.ai/next/reference/privacy-and-ai-use.md): What data sureops and Mintlify collect when you use this documentation site - [Helm chart deploy](https://docs.sureops.ai/next/self-hosting/helm-chart-deploy.md): Self-host sureops on your own Kubernetes cluster using the official Helm chart - [Analytics](https://docs.sureops.ai/next/using-sureops/analytics.md): Incident metrics, MTTR trends, agent performance, and cost - [Chat with ECHO](https://docs.sureops.ai/next/using-sureops/chat-with-echo.md): Ask your organization's AI assistant about incidents and system health - [Command Center](https://docs.sureops.ai/next/using-sureops/command-center.md): Your real-time incident operations dashboard - [Incident Hub](https://docs.sureops.ai/next/using-sureops/incident-hub.md): Deep-dive view — stages, HITL approvals, actions, and timeline - [Incidents](https://docs.sureops.ai/next/using-sureops/incidents.md): Viewing, filtering, and managing incidents in sureops - [Pre-mortems](https://docs.sureops.ai/next/using-sureops/pre-mortems.md): Proactive risk analysis before changes ship - [Problems](https://docs.sureops.ai/next/using-sureops/problems.md): Problem records, recurrence chains, and post-incident analysis - [Customer Landscape](https://docs.sureops.ai/next/your-rules/customer-landscape.md): Two paths to establishing your service topology — author .sureops/ yourself or accept a sureops-generated starter PR. - [The .sureops/ Repo Layout](https://docs.sureops.ai/next/your-rules/repo-layout.md): How to structure your service contract repository so sureops agents understand your environment. - [Runbooks & Knowledge Base](https://docs.sureops.ai/next/your-rules/runbooks-and-knowledge-base.md): How to author runbooks that agents use during incident diagnosis and resolution. - [services.yaml Reference](https://docs.sureops.ai/next/your-rules/services-yaml.md): Field-by-field reference for the sureops service catalog contract.