Blog

Cockroach Aegis: Redefining Database Observability and Operations with Agentic Intelligence

Published on October 7, 2026

0 minute read

    We've introduced Cockroach Aegis, an always-on system of intelligent agents for CockroachDB, available now in preview as part of Cockroach Continuum. Aegis continuously learns your cluster's workload, tracks every notable operational event, and builds a living memory of what normal looks like. It uses that memory to surface problems you'd otherwise miss, explain why they're happening, and tell you exactly what to do about them through recommendations.

    The database is only as scalable as the people who run itCopy Icon

    CockroachDB runs mission-critical workloads, the ones that can't lose data and can't go down. Cockroach Continuum takes that further by decoupling storage from compute and by implementing virtual clusters on a shared host cluster, both of which drop the cost floor low enough that an AI agent can spin up an isolated virtual cluster on demand and tear it down when it's done. A single deployment can now fan out into hundreds or thousands of virtual clusters. That kind of scale is hard to manage and troubleshoot by hand, and getting full value out of it can take deep CockroachDB expertise. The goal isn't to add more people every time the footprint grows. It's to let the team you have operate at far more scale than they could before.

    For operators it shows up as information overload. Something breaks at 2 a.m. and you're living across eight screens: the alert, three dashboards, the log tail, a couple of console pages, and a Slack thread asking if it's the database. Every tab tells you a sliver of the truth and none of them tells you the cause. Developers feel a quieter version. They aren't DBAs and shouldn't have to be, so a slow query investigation becomes a ticket that waits in a queue, and a fix that should take an afternoon takes a week. When there's no DBA at all, that gap lands straight on the developer.

    Agentic development is pouring fuel on this problem. AI coding assistants ship features and stand up workloads faster than any team can review them, which means more surface to watch, changing faster than anyone can watch it. The signals are almost always there in advance. Catching them just depends on the right person looking at the right moment and no human team can watch everything, all the time, and connect the dots fast enough.

    The teammate that never clocks outCopy Icon

    Think of Aegis as a new teammate that is an expert in monitoring and running Cockroach Continuum, and it works like a strong junior DBA hire: it watches your cluster around the clock, catches what your team would miss, explains why, and tells you what to do. Your only job is to review its work and decide when to act. Currently, Aegis monitors, learns, detects, diagnoses, and recommends and does not make changes to your cluster. Every recommendation comes with the context and remediation steps you need, and the action stays yours.

    It checks in so you don't have toCopy Icon

    Aegis does its ambient monitoring through automatic check-ins. It wakes on its own, observes the cluster, records what it found, and sleeps. The cadence adapts: healthy and quiet means it backs off to around every 30 minutes; something worth watching means it tightens to a few minutes. It can also register its own alert rules, for example unavailable ranges above zero, and wake early the moment a condition trips.

    Traditional monitoring runs on static thresholds which only fire once a metric has already crossed a line you set in advance. They don't adapt, they don't investigate, and they don't notice the slow drift that never quite trips a threshold. Aegis is the thing that's always looking, steering its own attention to the clusters and moments that need it, without anyone driving it.

    From a flag to a root causeCopy Icon

    When a check-in finds something off baseline, Aegis automatically opens an investigation: a structured record of it forming a hypothesis, gathering evidence, and confirming or ruling it out, and deciding on its own when it's reached a conclusion.

    Regular monitoring stops at the flag. It tells you p99 latency spiked and hands you a chart, and finding the cause is still all manual. Aegis does that chain for you, from the slow query to the plan to the stale statistics behind the bad estimate, connecting symptoms to root cause. Healthy clusters generate few investigations; troubled ones generate more, and each leaves a trail you can follow instead of a chart to decode.

    What Aegis hands youCopy Icon

    Check-ins produce two artifacts. Reports are Aegis's living memory: a current snapshot of cluster health, topology, and the workloads running on it. Recommendations are concrete actions tied to evidence, and they're the main thing to pay attention to. A recommendation might be to drop an unused index, refresh statistics, or look into a node that keeps restarting, and each one explains the justification, cites the metrics behind it, links to the docs, and gives the exact steps, often with runnable SQL.

    Most observability tooling gives you data, not decisions, and turning dashboards into "so what do I do" still falls on someone who knows the system. Aegis hands you the move, not just the reading. Recommendations are also deduplicated, evolved, and tracked over time, so it won't rediscover the same problem every check-in, and it surfaces the quiet ones that rarely page anyone but cost money and latency every day: unused indexes, stale statistics, missing indexes on hot queries, contention, configuration drift, and workloads that have outgrown their schema or cluster configurations.

    It gets smarter as your cluster changesCopy Icon

    Aegis learns your cluster instead of running off fixed rules. It starts building a baseline the moment it's activated, with pattern recognition inside the first hour, and refines that memory after every check-in as nodes, traffic, and applications change.

    Static rulebooks rot. A threshold that made sense at launch is wrong six months later, and teams end up muting the alerting that was supposed to protect them. Because Aegis adapts, its reports and recommendations stay accurate for the cluster you actually have. Because Aegis is a managed service by Cockroach Labs, intelligence improves automatically, drawing on years of experience on troubleshooting incidents and runbooks, and getting better automatically as Cockroach Continuum ships with new features.

    Ask Aegis anything about your clusterCopy Icon

    Aegis also supports cluster-aware Q&A. It lets you investigate alongside Aegis in plain language, with answers grounded in not only Cockroach Continuum documentation but also what it knows about your specific cluster, because chat sessions share the same memory and tools as a check-in.

    A general AI assistant gives you textbook answers with no idea which regions you run in or what changed last night. The alternative has been waiting on the one person who holds or knows how to retrieve that context. Aegis pulls in an expert into the loop the moment you have a question, whether you're mid-incident or planning a change.

    Bring your own agentCopy Icon

    Aegis has an agentic ecosystem built around a hosted MCP server and a CLI. Your own agents and tools can also pull its reports and recommendations directly, or delegate CockroachDB-specific work to Aegis and get back structured, cluster-aware context.

    Without this, every team rebuilds CockroachDB expertise inside every tool, and every agent that touches a cluster needs its own way in. Aegis gives all of them one safer path to live cluster context and becomes the specialist the other agents delegate to, whether that's a coding agent tuning queries before release or an SRE's agent pulling recent changes during an incident.

    Secure by design, not by afterthoughtCopy Icon

    Security shaped how Aegis connects, what it sees, and what leaves your environment. It's read-only, enforced in more than one place: Aegis parses every query with a CockroachDB SQL parser and runs it in a read-only transaction that gets rolled back, so any mutation is rejected regardless of privileges. Aegis also connects through a dedicated, identifiable SQL user with metadata-scoped grants, so its queries show up in your audit logs and you can see exactly what it ran. If you don't grant access to something, Aegis can't see it.

    Aegis has been designed to read system metadata, cluster shape, query fingerprints and latencies, metrics, event logs, settings, and schema names, never your application table data or row values. Every customer instance is fully isolated, access is protected by SSO, and while some parts of Aegis use AI, your data is never used to train or fine-tune any model.

    Get startedCopy Icon

    Aegis is available now in limited preview as part of Cockroach Continuum. Sign up here and point Aegis at a cluster for it to find the problems your team would miss, help troubleshoot the ones you're already staring at, and tell you exactly how to fix them. No deep CockroachDB expertise required.

    © 2026 Cockroach Labs. All rights reserved.
    Privacy
    Security