Creuto is now an OpenAI Select Partner Read More

Software Architecture & Technical

Automated incident remediation: where the human stays

Sentry, Datadog and Cloudflare now route production errors to coding agents. What automated incident remediation fixes, what it cannot, and the guardrails.

Automated incident remediation: where the human stays

Three observability vendors now ship the same loop: an error is detected, an agent is handed the stack trace and the code, and a pull request appears. None of them will merge it for you. That boundary — automated incident remediation that ends at a branch, never at production — is the one design decision that separates a useful integration from an outage you caused twice.

This post covers what has actually shipped and what is still beta, the permissions each integration asks for, which classes of bug the loop closes and which it cannot touch, and the guardrails worth insisting on before you turn it on. The short version: let the agent write code, never let it write to production, and keep a human on the merge button.

What has actually shipped, and what is still beta

As of 1 October 2026:

ProductWhat triggers the agentWhat it produces
Cloudflare Issues (Workers)Automations on occurrence thresholds or recurrence after a quiet periodSends the failure summary and diagnostic context to Claude Code, Cursor, Devin, a webhook or an incident tool
Sentry Seer AutofixAutomatic run when an issue has 10 or more events, occurred in the last 14 days and clears a fixability scoreRoot cause, a proposed solution, a code patch, and optionally a PR
Datadog Bits CodeAutomations on findings from Error Tracking, APM, Code Security, Test Optimization, or a scheduleA pull or merge request, optionally in draft mode, or a Slack notification

Cloudflare Issues is explicitly in open beta and free during the beta period. It needs no instrumentation: you set observability.issues.enabled to true in your Wrangler configuration and deploy, with Wrangler 4.134.0 or later. It records uncaught exceptions, failed invocations, 5xx responses and logged errors, and groups repeats so you can see when a failure first appeared and whether it is accelerating.

Cloudflare's own description of the loop is worth quoting because it is unusually honest about where it stops. The agent destinations are configured with real credentials — a routine ID and token for Claude Code, an automation webhook URL for Cursor, an API token and organisation ID for Devin — and then the human "review[s] the pull request, deploy[s] the fix, and mark[s] the issue resolved". The product ships the first 80% of the loop and declines the last 20% on purpose.

Sentry splits its status more finely. Autofix and Code Review are generally available; the Seer Agent is in open beta. Autofix runs four stages — root cause analysis, a proposed solution you can edit, code generation across potentially multiple repositories, and PR iteration against CI failures and review-bot feedback on GitHub. It can also hand the code-generation stage off to an external coding agent.

Datadog Bits Code is the broadest of the three, wired into APM, Error Tracking, Code Security, Cloud Cost, Containers, Continuous Profiler, Test Optimization and Sensitive Data Scanner, with automations available across Error Tracking, Test Optimization, APM Recommendations, Code Security and custom prompts. Its docs carry the sentence that should be in all three: "Bits Code never auto-merges PRs or MRs."

Automated incident remediation works for a narrow class of bug

The honest scope is smaller than the marketing, and it is defined by one property: the fix must be derivable from a stack trace plus the code, in one repository, without a product decision.

Bugs this closes well:

  • Null or undefined dereferences, unhandled rejections and missing guards — deterministic, single-frame, obvious in the trace.
  • Input validation gaps that a malformed payload walked straight through.
  • Breakage from a dependency API change, where the correct call is in the dependency's own docs.
  • Infrastructure-as-code misconfigurations, where the finding names the resource and the fix is a field.
  • Flaky tests, where the agent can run the suite repeatedly and verify stability.
  • Static analysis findings with a known fix shape — the same ground covered by agentic autofix on security findings.

Bugs it does not close, and where sending it anyway costs you time:

  • Saturation and capacity incidents. Nothing in the code is wrong. The fix is capacity, a limit, or shedding load.
  • Concurrency and ordering bugs. A single trace of a race condition describes one interleaving. A patch derived from it usually moves the race rather than removing it.
  • Multi-service incidents. Datadog states the limitation directly: Bits Code "does not support multi-repository investigations". Most real incidents in a distributed system cross a repository boundary.
  • Anything whose correct response is a rollback. The right action is reverting a deploy in minutes, not reasoning about a patch.
  • Data corruption and security incidents. These need containment and a forensic record before anything is changed.
  • Symptom-at-the-top traces. When the exception is three services downstream of the cause, the agent will produce a confident, well-written patch for the wrong file. This is the expensive failure mode, because the PR looks correct.

A reasonable rule: route automatically only where the error group is stable, frequent and scoped to code you own. Sentry's thresholds encode roughly this — 10 or more events, inside 14 days — and they are a sensible floor to copy even if your tooling does not enforce them.

Guardrails: scoped permissions, no production writes, a human on merge

Scope the repository permissions to what opening a PR requires. Datadog's setup guide is specific, and the list is a good template whatever you use: on GitHub, repository contents read and write, pull requests read and write, and the push event. CI iteration adds checks read and commit statuses read only. On GitLab, a service account with the Developer role and a token scoped to api, write_repository and read_user. On Azure DevOps, contribute, contribute to pull requests and create branch.

Note what is absent: no admin, no settings, no secrets read, no deploy. Note also the one escalation the docs flag — adding workflows:write to work around a GitHub branch-creation quirk "allows Bits AI to create workflows in your repository and has security implications". Granting write access to CI definitions means granting the ability to change what runs on merge. We would not grant it; if branch creation fails, fix the branch count.

Keep the agent out of production entirely. All three of these products write to git, not to your infrastructure, and that is the property to preserve when you wire up anything custom. An agent with a production credential is not an incident responder, it is an unreviewed deploy path. We take the same position we take on agents and security controls: if the control can be routed around, it is not a control. A read-only production role for diagnosis is fine. A write role is not.

Keep the blast radius physically small. Datadog runs each repository in its own isolated sandbox, with an organisation-level internet access policy governing outbound traffic after setup and an allowlist of package registry domains per language. Organisation secrets are available as environment variables only during environment setup, not during agent execution. That last distinction is the one most home-grown setups get wrong — they export the whole environment and call it convenience.

The docs also carry a warning worth reading twice: the permissions that manage organisation environment variables and secrets do not require source code access, yet they let a holder "influence environment setup and agent execution across all repositories in your organization". Agent configuration is a privileged surface. Treat it like production IAM, not like a dashboard setting.

Put the human on merge, and mean it. Sentry lets you cap automation at a stopping point regardless of the model's confidence: stop after root cause, stop after plan, or stop after a PR is drafted. PR creation can be disabled organisation-wide in advanced settings, which removes the create-PR button entirely. Datadog never auto-merges. Our default for client platforms is the middle setting — the agent may analyse and propose, and anything irreversible waits for a person.

What to log so you can reconstruct what the agent did

An agent that fixes production is a system you will eventually have to audit, so instrument it before you need to. Datadog's sessions model does most of this for you: each run captures the analysis, the actions and the resulting code changes, and sessions are shared across the organisation by default so a teammate can open one and see the reasoning. On the PR itself, Bits Code keeps a single status comment up to date with the run state and the link back to the session.

Where you build your own routing, the minimum record is: the alert that triggered it, the exact context handed to the agent, every command it ran, the diff it produced, who approved it, and what happened to the error rate afterwards. That last field is the one people skip and the only one that tells you whether the loop is working. The same instrumentation discipline we argue for in AI agent observability applies here — traces and cost caps on the agent itself, not just on the service it is fixing.

The strongest objection: a human on merge is a rubber stamp

The argument against keeping a human on the merge button is not that review is unnecessary. It is that review degrades. Give an on-call engineer twelve agent-authored PRs at 03:00 and approval becomes a reflex, which is worse than automation because it launders an unreviewed change through a process that looks reviewed. On that reading, a human gate is theatre with extra latency.

It is a real failure mode, and the answer is not more gates — it is making each PR cheap to judge. Require the PR to carry the reproduction, the failing test that now passes, and the agent's stated root cause, so the reviewer checks a claim rather than reconstructs one. Cap concurrent agent PRs per repository so the queue cannot outrun review. And route automatically only for the narrow bug classes above, so the reviewer's prior is "this is probably right" rather than "this is probably plausible". A gate nobody can pass attentively should be narrowed, not removed.

Turning it on without regretting it

  1. Start in draft mode. Datadog's automations can open PRs as drafts or post to Slack instead. Run a month that way and read the diffs before anything is mergeable.
  2. Pick one service and one error class. A Workers service with clean stack traces is a good first target — the failure modes that show up early in Workers are mostly the deterministic kind this loop handles.
  3. Grant the permission list above, and nothing else. Review the grant quarterly, including who can edit agent environment variables.
  4. Measure one number. Of the PRs the agent opened, what share merged without a human rewriting the fix? Below roughly half, the routing is too broad — tighten the trigger, not the review.

The uncomfortable part of this pattern is that it does not reduce the work of incident response so much as move it. You spend less time writing the patch and more time deciding what the agent is allowed to touch, which is harder and does not feel like progress. That decision is the engagement — most of what we do in DevOps and cloud engineering around agents is drawing that boundary and then proving it holds. Draw it once, write it down, and the loop is worth having. Skip it and you have automated the production of plausible patches.

Frequently asked questions

An agent can open a pull request automatically, but none of the shipped integrations deploy it. Cloudflare, Sentry and Datadog all end the loop at a branch: a human reviews, merges and deploys. Datadog's docs state plainly that Bits Code never auto-merges pull or merge requests.

It is safe when the agent writes only to git, holds repository permissions limited to contents and pull requests, runs in an isolated sandbox with outbound network restrictions, and has no production credentials. Automated incident remediation becomes unsafe the moment the agent can change infrastructure directly.

Anything irreversible: merging, deploying, rolling back, changing data, and containment during a security incident. Capacity decisions and multi-service diagnosis also stay human, because a single stack trace cannot support them. The agent proposes a patch; a person decides whether it ships.

Deterministic single-repository failures: null dereferences, unhandled rejections, missing input validation, dependency API changes, IaC misconfigurations, flaky tests and static analysis findings with a known fix shape. Race conditions, saturation incidents and cross-service failures fall outside what a trace plus one repository can explain.

Grant repository contents and pull request read-write plus push events, and nothing more. Avoid workflow write access, keep secrets available only during environment setup rather than agent execution, restrict outbound network traffic to package registries, and review who can edit agent environment variables.

Sentry runs Autofix automatically when an issue has 10 or more events, occurred within the last 14 days, and clears a fixability score from Sentry's model. You can cap how far it goes — stop after root cause, stop after a plan, or allow a drafted pull request.

Record the triggering alert, the exact context given to the agent, every command it ran, the diff produced, who approved the merge, and the error rate afterwards. The last item is the one teams omit, and it is the only one that shows whether the loop is working.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

1 Oct 2026

·

10 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved