Sign up
LearnPublished October 5, 2026

Inside Slipstream’s Recall Pattern

How Slipstream assigns six types of engineering maintenance work to agents instead of engineers.

Simon Fletcher

Simon Fletcher

Inside Slipstream’s Recall Pattern

Our previous Slipstream article introduced a recall pattern that addresses issues without human intervention. This article explains that pattern in detail.

Engineering teams have long spent time on six distinct types of maintenance work (outlined below). Slipstream now assigns these six types of maintenance work to agents, rather than humans, by default.

We explain what each agent does, how it validates its work, and where humans must remain involved.

The recall pattern consists of six agents

Slipstream's recall pattern detects issues before humans do. The pattern consists of six separate agents, each responsible for monitoring a specific type of maintenance task.

The 6 agents are:

  1. Bug triage
  2. Production error handling
  3. Security patching
  4. Codebase health improvements
  5. Stale feature-flag cleanup
  6. Rollout monitoring

The maintenance work handled by these agents already existed: unresolved bugs, recurring production errors, security vulnerabilities, and codebase degradation.

Historically, limited engineering resources left much of this work unresolved unless it became critical. Slipstream assigns ownership of these tasks to agents so it gets handled much faster.

Each agent uses the same basic process:

  • A human, observability tool, or audit opens a GitHub issue.
  • An agent picks up the open issue and first verifies that the problem exists.
  • If the agent cannot reproduce the issue, it does not attempt a fix.
  • If the agent confirms the problem, it implements a solution.
  • The agent opens a pull request and monitors it until the work meets Slipstream's criteria.
  • Humans review and merge the changes. Some agents are allowed to merge after a set deadline, if there is no human objection.

Two rules apply across all workflows. First, an agent must verify that a problem exists before attempting a fix. This requirement prevents agents from spending time on unvalidated issues. Safeguards and feedback systems also keep agents focused on work that matters.

What each agent does

Each agent owns a specific maintenance task. Its goal is to handle that task the way a software engineer would, and ship the change. The sections below describe the workflow for each task.

1. Bug triage

Humans continue to file bug reports as usual. Each day, agents receive reports that have not been addressed. The agent first determines whether it can prove the bug by confirming the root cause, reproducing the problem, and determining whether the fix is minor. 

If it can, the agent creates and commits a failing test, applies the fix, and verifies that the test passes. 

If it cannot, the agent reports its findings, including a potential proposed fix and any unresolved questions.

A pull request with a validated fix merges automatically after seven days unless a human intervenes. A review, hold, or comment pauses the process and gives humans time to respond.

2. Production error handling

Agents audit Datadog, our engineering observability platform, every day for new production errors. Most errors are expected results of user actions or state changes rather than true bugs. The agent primarily triages these errors by assessing their effect on users and suppressing non-critical issues. 

It creates GitHub issues only for errors it considers serious, and agents cannot close issues they have rated as serious. When the severity is uncertain, the agent takes the cautious approach and rounds up the severity.

After an issue enters the backlog, a different agent generates the fix. Keeping triage and resolution separate creates a clear audit trail. Separate runtime monitoring and on-call processes handle more widespread impact. This agent focuses on smaller, papercut-style runtime errors.

3. Security patching

Daily audits identify potential vulnerabilities. The security agent bears the strongest burden of proof. Before implementing a fix, the agent must demonstrate that the exploit works. It documents the successful attack, the fix it applies, and the failure of the same attack after the fix. 

If the agent cannot reproduce the vulnerability, it does not attempt a fix. Instead, it flags likely false positives, notifies the engineering team, and leaves the issue open. 

Only humans can close issues or merge fixes by this agent.

4. Codebase health improvements

Each week, an agent runs Fallow, our static analysis tool that flags dead code, security issues, and style violations against the codebase. We use Fallow for static analysis so agents receive consistent backpressure type feedback. A low score creates a GitHub issue. 

The agent then refactors the code while preserving its behavior, which the existing test suite verifies. The goal is to improve the Fallow score. These pull requests do not close their own issues because only the next audit can confirm that the problem is resolved. 

Only humans can close issues or merge fixes by this agent.

5. Stale feature-flag cleanup

A feature flag is considered stale after it remains fully enabled or disabled in production for two weeks. An audit finds these flags and opens tracking issues. An agent then removes the flag and its related code paths. 

As with bug fixes, these pull requests merge automatically after seven days unless a human intervenes.

If someone closes a pull request without merging it, the agent treats that action as a signal that the feature flag is still needed.

6. Rollout monitoring

Every production deployment starts a watcher. The agent reviews the changes and monitors relevant telemetry for ten minutes. When it can connect a regression to a specific change, the agent initiates a revert and notifies the responsible person in Slack.

If that person does not respond, the agent escalates through team channels and direct messages. After 30 minutes without a response, the agent automatically reverts the change. 

When the cause is unclear, the agent reports what it found and takes no further action. The rollout agent detects and mitigates problematic deployments, supports autonomous agent changes, and protects the user experience.

Integrating feedback into engineers' existing workflow

These agents depend on engineering feedback to show whether their work is correct. Building the agents themselves was the easier part.

The system works only when engineers can provide feedback without adding extra work. Moving from backlog-driven maintenance to automated fixes happens gradually. Engineers need to be able to review agent-generated work as easily as they review work from other engineers.

The proof-before-fix rule is critical to maintaining that review process. If an agent starts submitting work that engineers do not need to review, engineers lose trust in the queue and stop providing feedback. 

Without real feedback, an agent spends tokens without improving.

The agents operate directly in GitHub, where engineers already work. GitHub issues track all activity, and engineers review agent pull requests through the same process they use for human contributions. Reviewing an agent's pull request requires no new review workflow.

This integration also addresses a problem with agents that generate work by default. Over the past nine months, we have learned that large language models (LLMs) produce output regardless of whether the output is useful.

The proof-before-fix rule is the first safeguard for each agent, and human review in GitHub provides another. Engineer feedback gives the agent information it can use to improve, handle harder cases, and produce more useful results.

How agents earn the autonomy to run unattended

An agent earns autonomy by proving that it can act reliably. Autonomy is not a setting that we turn on.

Every agent starts with human approval required for each action. The agent builds a track record over time. Once that record meets the standard we have set, the agent graduates to greater autonomy.

The agent can then act on its own.

Graduation does not remove the agent’s guardrails. So far, graduation has changed only one requirement: human approval becomes optional instead of required.

Each agent keeps a record of its previous mistakes. Before applying a fix, the agent reviews those errors. When a mistake provides a broader lesson, the agent adds the lesson to the record. This record is agent memory, and it helps the agent become more resilient over time.

Engineering time saved in the last 30 days

The six agents have saved approximately ten hours of engineering time per day, based on the last 30 days of GitHub data.

Rollouts and feature flags record timestamps for each issue, so we can measure their time directly. Production errors, security, codebase health, and bugs do not have the same tracking. For those four agents, we estimate time saved by multiplying the number of issues by the average time an engineer spends on each issue.

Measured directly

Rollout monitoring covers approximately 25 production deployments each day, with a median observation window of 12 minutes. Human engineers did not previously monitor every deployment consistently, which was one reason we created this agent. The rollout agent now accounts for 3.5 to 7.5 hours of deployment monitoring per day.

Feature flag removals average 1.3 pull requests per day. An issue takes an average of 18 minutes to move from open to pull request to ready. Manual flag removal typically requires 45 to 60 minutes, including review follow-up. The feature flag agent saves approximately one hour per day.

Estimated from throughput

Production monitoring automatically files error reports, and the agent resolves nearly all of them within minutes without human review. Only two reports this month were rated high or critical. We estimate that this saves a few hours of engineering time per day.

This month's security audit findings produced 27 fix pull requests through the security agent. The agent also discarded five false positives before any code was written. We estimate that this saves a few hours of engineering time per day.

Too early to measure reliably

The bug and codebase health agents are still too new for reliable estimates. Together, we currently estimate that they save less than one hour per day. We expect that figure to change as both agents mature.

Rollout3.5 h/day
Feature flags1.0 h/day
Production errors2.5 h/day
Security2.5 h/day
Codebase health + bugs0.7 h/day
Total~10 h/day

The total is approximately one and a half to two engineers' worth of daily maintenance work. This represents the minimum impact the agents currently provide.

“AI has increased our development velocity by an order of magnitude. But as in physics, the faster we go, the more drag we create. Every leap in output adds surface area, complexity, maintenance, and vulnerabilities. Reliable autofix loops aren’t a nice-to-have; they keep that drag from becoming our speed limit, letting us move fast without being weighed down by everything we’ve already shipped.”

Ryan Daigle

Ryan Daigle

VP of Engineering

Autonomy depends on meeting a quality standard

Every agent starts with human review. Whether an agent continues to require that oversight or earns autonomous merge rights depends on the quality and consistency of its results over time.

We evaluate that performance qualitatively and quantitatively. GitHub feedback shows whether engineers continue to find problems or whether feedback has stabilized. Consistent approval with little feedback indicates that additional human sign-offs may be slowing changes without improving safety.

Bugs, feature flags, and rollouts followed this progression. Each agent first operated with full human review and earned merge rights only after receiving consistently positive feedback.

Security, production errors, and codebase health have not reached the same level of autonomy yet. We are working toward that standard. These agents have a higher threshold to hit because mistakes in these areas carry greater costs.

The agents handle maintenance so engineers can build

Operating GlideOS has always required attention to bug triage, flag cleanup, security patching, and code health. Traditionally, engineers had to reserve part of their time for this maintenance instead of using that time to create new customer value.

Slipstream now assigns this maintenance work to agents. Engineers can focus on tasks that require their expertise, while agents handle issues they can validate and resolve independently. The more deterministic work that can be validated by guardrails effectively. 

Engineers remain responsible for deciding what to build next and reviewing work returned by the agents. With less maintenance work to review, they can spend more time shipping features.

The compounded output effect is revealed in our changelog, where Glide’s engineers are now shipping new features almost daily, rather than monthly.

Create a Glide app in five minutes, for free.

Sign Up
Simon Fletcher
Simon Fletcher

Principal Software Engineer at Glide

Glide turns spreadsheets into beautiful, intelligent apps.