Between a Wednesday evening and a Friday afternoon in April, MOLT, the tool Cockroach Labs ships for migrating databases to CockroachDB, gained the ability to migrate from IBM Db2 sources.
Db2, one of the first commercially available relational databases, has a rich SQL dialect, a complex type system, and a native wire protocol. Supporting it in MOLT meant a new schema converter, a new Fetch path for pulling rows out and getting them into CockroachDB, a new Verify path for comparing them once they were loaded, a vendored ANTLR grammar, a Docker image for CI, and more than ten thousand lines of test fixtures. When we added Oracle support to the MOLT tools in 2024, the equivalent work took 9 months, and cost around $160K in engineering time.
Db2 took less than two days, and no human wrote any of the code. The token bill was $4172: 164x faster, and 38x cheaper.
It all started with a single GitHub issue describing the ask:
A planning agent then read it, decided it was too large to treat all at once, and decomposed it into fifteen sub-issues with an explicit dependency graph: first the foundation, then the type system, then the row iterator, then Fetch, then Verify, then Convert, then CI, then test data. Two of the sub-issues were judged as too large in their own workups and were decomposed again. Along the way the agents filed another dozen issues against their own work: fixture gaps, a type-mapping bug, and an isolation-level fix. By the time the parent issue closed, thirty-two sub-issues had been opened under it, twenty-seven pull requests had merged, reviewers had sent work back fifty-five times, and nine issues had been escalated, two of them to a human. Test coverage, as measured against PostgreSQL (our best-tested dialect) was equivalent.
But how did all of this happen? What was driving this, and more importantly, what was ensuring that what came out the other side was high quality? Before we opened the issue for Db2 support, we first built a coding hospital which we call MOLT Sinai. MOLT is the name of CockroachDB’s migration tool set, and once we had decided to model the pipeline on a hospital, MOLT Sinai was too good of a portmanteau to resist. That’s because Mount Sinai is a prominent teaching hospital in both Toronto and New York - the cities in which we live. The name stuck, and so did the vocabulary. Issues are patients. Merging is discharge. The humans in charge are the Chiefs of Medicine.
Agents writing a lot of code quickly is not unique to this story, as software factories are popping up all over the industry (like this one at OpenAI). What’s unique to our case is what stood between those agents and main, why we built it to look like a hospital, and what happened when we ran it on real work for five months.
Why a teaching hospital
The problem with coding agents is that although they’re more than capable of writing code, it’s hard to trust them to produce production-grade code that is well-tested, maintainable, and within the desired scope. Agents are relatively cheap to run, but that's little comfort when their code is going to land in customers' hands. We'd much rather have the agents take longer to run (even 10x longer) than land the wrong data in a customer migration to CockroachDB. Speed matters, but quality matters much much more.
Several attempts to automate agent coding have appeared over the last year, many of which have been optimized for throughput. Gas Town, for example, runs many agents in parallel against a codebase with a merge-handling layer to sort out collisions. That is a good design when the bottleneck is how fast code gets written. And there’s this experiment from Cursor in which they rebuilt SQLite so quickly that they needed to write their own version control system to handle the change volume. It’s an interesting experiment, but clearly it’s not designed to ship quality database code to users. In the vast majority of early agent automation experiments, speed was the main motivator.
Our motivation was different. We build a distributed SQL database, and a suite of tools to migrate workloads to that database. A subtle bug in a migration tool could corrupt someone's data on the way into a database that is supposed to maintain correctness above all else. As a result, the question we cared about was not "how many PRs per hour?" but rather, "how many of those PRs would we have been embarrassed to have merged ourselves?" With that reframing it was clear that a throughput optimized swarm of agents wasn’t a good model.
So we looked for a model with long experience managing high-stakes work done by staff of uneven ability, with mandatory handoffs, mandatory second opinions, and literature on what goes wrong when the handoffs are sloppy. And we landed on a teaching hospital.
The metaphor gave us these roles:
Around that core sit the other actors you would expect to see if you’ve ever been inside a hospital (or have watched The Pitt): a Charge Nurse who rounds every thirty minutes looking for stuck patients, an Infection Control agent that can lock down every workflow in the building if main breaks, a Safety Department that writes weekly process reviews and runs post-incident conferences, and a Research Department that proposes new work.
We don’t claim that this is cheaper than a swarm or simpler than a single agent. It does however produce a higher-quality product than either. When you’re in the database business, you’re prepared to pay for that quality. As you’ll see below, it also brings some of the regrettable aspects of teaching hospitals (bureaucracy, wait times, and cost) but we believe that those are worthwhile tradeoffs.
How it’s built
Everything in the hospital runs on GitHub Actions. The state of every patient lives in the labels on its issue, as well as in some scratch files that the hospital creates to keep track of each patient’s progress.
Each stage is a GitHub workflow that starts when its “label in” lands on the issue. When the agent finishes, it writes a “label out”, and that label is the next stage's “label in”. Dashed arrows send work back: a rejected plan to workup, a red CI run or requested changes to treatment, and a stuck stage up to the Chief.
Not shown: four departments run outside this flow on their own schedules. The Charge Nurse re-runs stalled stages, Infection Control locks the hospital down when main breaks, and the Research and Safety departments file new issues and policy PRs, which enter through the Chief like any other patient.
Applying a label fires a GitHub workflow which loads a role-specific prompt so that the agent acts as an employee in the hospital. Once it’s completed its work, the agent adds a structured note to the file containing that issue’s chart, and applies the next label in the flow. The note carries a sentinel like <!-- SINAI:TREATMENT_PLAN --> so that later agents can find and parse it.
A few basic hospital rules, encoded in the skills, provide most of the safety:
No code before a plan, no plan without review. Before the Fellow fixes anything, it has to reproduce the problem, perform a differential diagnosis (to use hospital terminology), run whatever targeted investigation narrows it, and post a treatment plan: diagnosis, files to change, tests to add, risks, open questions. A Review Attending, a different agent instance running a prompt built around finding fault rather than fixing, reads the plan and approves, or rejects it. Only then does treatment start. We added this gate because an agent that confidently implements the wrong fix wastes far more time and tokens than one that stops to check its plan.
Don't improvise. If scope changes during treatment, the Fellow stops, posts a revised plan, and goes back to plan review. The instruction in the skill file is two words: "Don't improvise."
No shortcuts, especially around testing. An agent may not disable, skip, weaken, or modify a test to make it pass. If a test fails, the test is correct until proven otherwise. If the agent genuinely believes the test is wrong, the change to the test has to be its own commit, justified on its own merits and reviewed separately from the fix. Similarly, a passing test that does not actually exercise the change is worse than no test at all. All code reviews start by disabling the change and ensuring that all new tests fail. If they don’t they’re either out of scope, or incorrect.
Never flail. When a Fellow is stuck, it does not try harder. It writes an I-PASS (Illness severity, Patient summary, Action list, Situation awareness, Synthesis) handoff, which is a structured format borrowed directly from teaching hospitals, and then escalates to an Attending. The receiving agent has to write back what it understood of the handoff before it starts working. Studies of I-PASS in hospitals report a 50% reduction in adverse events during shift changes if physicians and nurses write down patient information in a standardized way. Our agents follow the same protocol.
The Discharge Nurse audits the reviewer. The last gate before merge does not re-review the code. It verifies that the review happened properly: an approval exists, the review template was filled out, no threads are unresolved, CI is green, the commit history is clean. It posts a checklist showing exactly what it verified, and if any box cannot be checked, discharge fails with that box left unchecked and annotated. The rule exists because when you’re turning over control to agents, you want to ensure that they’re doing what they’re told, and sometimes they aren’t.
Check the precedents before bothering the Chief. Every decision the human Chief makes that generalizes, is recorded in an append-only precedent log. An agent that wants to escalate has to read the log first, and if a precedent applies, it follows the precedent and cites it instead of escalating. Only a human can create precedents. Over five months, this has cut down how often agents escalate to the Chief.
What it shipped
When we started working on MOLT Sinai, we created a mirror of the repo in which the MOLT tools are built. This mirror then became a proving ground where we could test the efficacy of the hospital model. Here’s what transpired in that repo between April 21 and September 11:
In just under 5 months, the hospital pipeline landed well over a million lines of code, split into changes small enough that they could be confidently reviewed by agents.
You may have noticed one curious number in the table above: 1299. Nearly half the issues filed in the repository were created by the hospital for the hospital: decomposition children, follow-ups the Fellow scoped out of a plan, proposals from the Research Department, corrective actions from the Safety Department. The pipeline generates its own backlog. A human decides what gets admitted, but the humans stopped being the main source of work in the second month.
While Db2 support was the first large feature we put through the hospital, that was just an experiment. After its first week of autonomous operation, we turned on “human approval mode”, which requires every merge to have a set of human eyes pass over it. It was only in this mode that we could be certain that the code was high enough quality to ship to customers. With the hospital having this extra safety layer, we went to work on what we really created the hospital for: building the Migration Assistant, an AI based tool which walks users through a complete migration from Postgres to CockroachDB, including schema conversion, data load and verification, and routine conversion. The Migration Assistant is in preview now for Postgres migrations, and was almost entirely written by the hospital workers.
What it cost
Since we started MOLT Sinai running in April, it’s consumed just over $135,000 in Claude tokens, which works out to around $84 for the average patient (i.e. GitHub issue). Issues typically spend one to two days in the hospital, often dominated by the time spent waiting on a human reviewer. In terms of LLM models, the hospital has evolved as models have advanced.
At first it was running an Opus model for every stage of the pipeline, but with the introduction of Fable, it now plans on Fable, and the Fable-generated plan chooses the model to leverage for the implementation. In cases where the implementation struggles on a Sonnet or Opus model, it can escalate to leverage a Fable model. This helps to keep costs down, while still allowing the more expensive (and capable) models to be leveraged when necessary.
As the chart above shows, with each successive model switch there’s some work required to tune agent skills to optimize cost. We perform successive optimization rounds every couple of weeks, which is one of the reasons the chart above is choppy. There are other reasons, such as the mix of work in a given period (UI tasks cost more because the agent has to process screenshots and other visual data) and quirks in our CI environment, but they’re too detailed to get into in an introductory blog post.
The hospital built itself
When we created MOLT Sinai, we built the initial set of workflows, skills and prompts directly into the mirrored repository. Soon after, we realized the value in what we’d built and suspected that others at the company would want to leverage it. To make this possible, we split the “Sinai” code (the actual hospital pipeline) out from the MOLT Sinai repo (the hospital pipeline’s first consumer) into a separate repo that others could embed into their repo if they wanted to leverage the hospital model. Today, the Sinai repository has about 1,570 commits. Roughly 1,340 of them – 85% – were authored by the pipeline itself.
We didn't plan this. The framework had bugs, and we sent them through the hospital like any other issue. The Charge Nurse workflow that does rounds every thirty minutes exists because an agent hit a race in GitHub Actions label events, filed an issue about it, and another agent designed and shipped the fix. It was also the most expedient way to evolve the hospital model quickly, and with almost every patient discharged, the hospital model has become more robust, efficient, and easier to use.
Today the Sinai model is in place in four repos at Cockroach Labs, with several other repos planning on adopting it in the coming months.
The instructions became legacy code
Each agent role is defined by a skill file: a Markdown document of procedure, rules, templates, and forbidden behaviors. They were written the way knowledge bases often are, by adding a rule every time something went wrong.
In July we audited them. Twenty-five skill files, a little over 100,000 words. About 23,000 of those words, 23%, were flagged as removable without changing a single gate, command, template, or sentinel. The dominant pattern was the same rule restated in several places: six thousand words of it across seventy findings. One file was forty-seven percent redundant.
The redundancy was costing us real money and time. The hospital workers load several skills at once. The Discharge Nurse loads about 21,800 words of instructions, roughly 29,000 tokens, on every single run, before it has read a line of the PR it is supposed to be checking. The base hospital-protocol skill is loaded by all fifteen agents, so redundancy there costs the most.
We suspect this is not specific to MOLT Sinai. Our prompts rotted the way code does: we kept adding rules and never removing them, until the files were too long for anyone to read them end to end. We only noticed because of the token bill. We do not have a linter for this yet, but we’ll likely be creating one soon.
Lessons learned
Five months of working with MOLT Sinai (and the Sinai model) have taught us a lot. Here are some of the more interesting bits.
What worked
The model does more than expected. Each role's skill file starts with a few sentences framing the role, and we’ve found those few sentences to be the most valuable. The Fellow's begins: "You are a Fellow conducting a diagnostic workup on an admitted patient (GitHub issue). Your job is to investigate thoroughly and produce a written treatment plan that another Fellow could execute." The Review Attending's starts: "You are a DIFFERENT agent instance from the Fellow who treated this case. Your job is to find bugs, security issues, regression risks, and deviations from the approved plan." This lays the groundwork as the underlying LLM knows very well how a teaching hospital works, what a differential diagnosis is, why a handoff has a structure, and what it means to be a second opinion. When you tell an LLM a persona to embody, it does it very very well. As an example, a Review Attending found a one-line docstring defect on a PR which was already several review rounds deep and declined to send it back for rewriting it on its own authority, writing that it was "surfacing it for your adjudication rather than unilaterally sending the case back to a Fellow over a docstring-only fix while the patient is exhausted." In another case an Attending deciding whether to re-decompose an issue the Chief had already ruled on wrote: "A Chief-owned decision whose premise the outcome disproved is returned to its owner, not improvised around." None of these sentences appear directly in a skill file. Instead, the agents wrote them because they were playing a part.
Plan review before code. Plans that get reviewed are rejected less than 10% of the time, but when they are, it’s the highest value review in the system, as it shifts the rework to the earliest point possible. In one example, the plan review of a telemetry fix to remove a customer’s hostname and database name from our captured logs uncovered a small gap where the information could still be exposed. If that hadn’t been caught in the plan review, it’s highly unlikely that it would have been caught in the code review, and the issue would have either shipped, or would have required a second issue to resolve. Catching bugs when planning saves mountains of future work.
Decomposition, with a hard ceiling. If a workup estimates that the change will exceed 1,000 lines of code (including tests), the Fellow must decompose it into sub-issues with a dependency graph. Any sub-issue that would itself exceed 1,000 lines is decomposed again. The point of the ceiling is to keep each change small enough that a reviewer, agent or human, can hold all of it in their head. The Db2 sprint resulted in 32 issues from a single admission. Tens of thousands of lines of code can not be convincingly reviewed by a single agent (or human) so the hospital makes the work easier by splitting it up.
The records. Every hospital issue records everything that it produced: the workup plan; the reproduction test the Fellow wrote before touching code, and the transcript of it failing; the treatment plan and each revision of it; every I-PASS handoff; every review verdict with its reasoning; a clinical note for every stage with the model, the tokens, the duration, and the cost. Records live in a companion archive branch that merges alongside the code so that the PR diff stays code-only, and is therefore easier to review. Five months later, we can easily reconstruct why any of 1,238 PRs was merged, using the issue thread alone, and the Safety Department's weekly reviews are possible only because that trail exists. Meticulous record keeping turned out to be one of the most powerful things that the hospital metaphor gave us.
What we changed, or still need to
The model sometimes does too much. The same prompting that makes the roles work also makes them bureaucratic. A Fellow's plan for one issue was simply "close as duplicate; PR #691 already shipped the fix." The plan review verified that this was correct but escalated anyway, because "the review-attending-plan-protocol's three documented outcomes (APPROVE / REJECT / ESCALATE) do not include 'close as duplicate',". A confirmed duplicate needed a human (something we’ve since resolved). A one-line UI copy edit, changing "Validating" to "Verifying" in a subtitle, went five rework rounds and reached the Chief over whether a low-risk PR could merge before a flaky CI run was green (the agents often try to merge with CI red, arguing that it’s not their issue causing it). A code review blocked a PR with no code defect at all, on the grounds that the PR description had a minor typo. And when a reviewer left non-blocking nits, the pipeline had a tendency to open new issues to track them (also since fixed), turning a minor comment into a new patient. Some of this we have addressed, but some of it is the price of a system that follows its own rules and puts quality as its #1 priority.
Reviewing needed a circuit-breaker. Issue 1777 was marked urgent and the fix was one line: call an existing function in one code path, the same way three sibling paths already did. It took eleven rework rounds over two days, first due to comment hygiene, then a false claim in the PR description, then a minor deviation from its original plan, then test quality, two rebases, two failed workflows, and a human-approval gate. Anything that takes more than three rework rounds gets a Safety review, and in one Safety review window, three complex issues went nine, nine, and seven rounds and consumed about $631 of the $1,046 the hospital spent that window. That same safety review documented that the code review process had no "countable non-convergence circuit-breaker," so non-converging loops "escalate late or never." The fix it proposed merged recently - the reviewer must now hand off to an Attending once a PR has gone four consecutive rework rounds that each surfaced a new blocking finding, or six substantive rounds of any kind. While this merged rather recently, and we don’t have good data to point to its efficacy, early signs seem positive.
The hospital is prolific, and not always in a good way. The hospital is prolific at finding work for itself. As noted above, nearly half of all issues were filed by agents. The Research Department, which surveys the codebase and proposes new features, originally ran every three hours and the volume of issues it created was so high that we couldn’t keep up with reviewing them for admission. We first throttled it to daily, and then we capped the number of its proposals that could wait unadmitted at ten. The problem is the issues it finds are largely good. It’s not infrequent where we’re working on a problem in the codebase with an interactive Claude session and it points us to the identical issue filed by the hospital weeks (if not months) earlier. That being said, pushing all of them through would be expensive and time consuming, and of dubious value. Finding the nuggets of gold in the river continues to be the hard part, and one that we haven’t yet solved.
We built a teaching hospital that doesn’t really teach. MOLT Sinai is exactly what you’d expect to see built by two senior engineers (with over 40 years of collective industry experience). The model is incredibly powerful, and can churn through issues while we sleep, only to have them quickly reviewed in the morning. Then, we can quickly load up the hospital with more issues to get resolved while we spend our days in meetings, reviewing them and discharging them before we leave for the day. A full day of meetings can now lead to 10 high quality merged PRs. But what about the new-grad who doesn’t have a full calendar, the software engineering judgment that decades of experience buys, and has a desire to deeply learn how to be a stronger engineer? How does the model work for them? In short, it doesn’t fully meet their needs. We’re currently working on enhancements to the hospital which will allow humans to take triaged issues and craft their own plans and have them reviewed by the hospital agents. As part of that review, they’ll be asked to defend their decisions, and challenged to think of other options that may yield more comprehensive solutions. Similarly, when presented with a reviewed plan, the human may choose to write the code (either by hand, for small fixes, or with the help of an interactive LLM-driven coding session) so that they can learn how to navigate the codebase, make implementation tradeoffs in real-time, and be measured on their performance in the resultant code review. We need to build the software engineers of the future, and regardless of what the agentic coding world looks like a decade from now, we’ll still need people with sound judgement and critical thinking skills..
Next Steps
We’re constantly working to improve the efficiency, speed and accuracy of the agents working in the hospital. Additionally, we’re working to make the hospital more engaging, educational, and skill building for our more junior engineers. Oh, and we’re also working on confirming that the IBM Db2 support is correct, as we weren’t reviewing any of the code that the hospital wrote during that initial test. Finally, we’re conducting a series of experiments to see what it would take to put the hospital back into fully-autonomous mode, at least for some classes of issues. If you’re interested in helping us solve one or more of the problems above, we’d love to hear from you.








