Bug Buster · Operational control

14 minute readReview draft

A bug report is evidence, not permission to change production

AI can classify reports, review pull requests, correct lint failures, reproduce faults, and prepare fixes. Bug Buster escalates regressions and risks introduced by a change instead of treating a clean diff as permission to release.

The operating problem

The queue is not slow because nobody can type the fix

A production report usually begins incomplete: “booking failed”, “the total looks wrong”, or “I cannot sign in”. The costly work is turning that observation into a bounded claim. Which user state produced it? Is it repeatable? Did data become incorrect or did the interface only display it incorrectly? Which products share the dependency? Is there already an incident, a duplicate, or a known limitation?

AI can reduce the search cost across reports, logs, traces, tests, and code. That makes it valuable early in the process, before a line is changed. It can assemble context and test a hypothesis quickly. What it cannot infer from a ticket is the authority to decide that the proposed behaviour is correct for the organisation.

Maintenance accelerates when evidence moves faster—not when production authority becomes ambiguous.

Decision path

A signal earns investigation before it earns a change

01 · decision

Report or PR signal

Capture the observed outcome, failed check, changed lines, and relevant state.

02 · decision

Reproduce or verify

Create a failing example, rerun the exact rule, or mark the uncertainty explicitly.

03 · decision

Bound impact

Identify affected systems, data, authority, and reversibility.

04 · evidence

Prepare evidence

A proposed change carries tests, scope, and a rollback path.

05 · refused

Release authority

The product owner decides whether the evidence is sufficient for this risk.

Each step changes the status of the claim. Observation, reproduction, impact, and proposed repair are evidence; release remains a separate decision owned by the product.

The plugin model

Bring product context to one triage discipline

Bug Buster operates as a common maintenance workflow paired with a lightweight product integration installed as aversioned dependency in each system (such as Garage CRM or Grand Total). The product-side plugin supplies bounded context: release identity, feature area, safe diagnostics, ownership, test commands, and links to the applicable runbook. The central service groups related reports, proposes severity, traces the affected surface, and prepares a reviewable change where the evidence is sufficient.

On an open pull request, the integration can run approved lint and test commands, connect failures to changed lines, and prepare mechanical corrections. New invariants, error paths, permission changes, migration hazards, and dependency risks stay attached to the pull request and are escalated to its reviewers rather than hidden inside a broader fix.

The plugin is not a permanent administrative tunnel. It should expose the least authority required for observation and preparation, with explicit capability scopes and an auditable request trail. Secrets, personal data, and unrestricted production access do not become acceptable merely because the consumer is an internal automation.

StageAutomation mayProduct owner retains
IntakeDeduplicate reports, inspect pull-request signals, and request missing contextDefinition of material impact and escalation obligations
PR preflightRun approved lint, test, and static-analysis commands; prepare mechanical correctionsDecision on findings introduced by the change and whether the pull request may proceed
ReproductionBuild a failing example in an isolated environmentJudgement about whether the example reflects intended behaviour
PrioritisationEstimate scope, affected versions, and evidence confidenceTrade-off between user harm, operational risk, and planned work
RepairPrepare a small change, tests, release notes, and rollback evidenceApproval of behaviour, migration, and release timing
LearningCluster recurring fault patterns and propose guardrailsDecision to change architecture, policy, or product design

Pull-request guard

Fix the lint error. Escalate the behaviour change.

Deterministic lint failures such as formatting, import order, or an unused symbol can usually be corrected and verified by rerunning the same rule. A failing test may also have a narrow correction when the intended behaviour is already explicit.

A pull request can expose a larger issue: an omitted authority check, a retry that duplicates an external action, an incompatible migration, or a test changed to accept the wrong domain behaviour. Bug Buster surfaces the introduced risk, identifies the affected boundary, and stops the automated repair path. Rewriting it silently would remove evidence the reviewers need.

Pull-request findingBug Buster responseEscalation condition
Deterministic lint failurePrepare the smallest correction and rerun the exact lint ruleThe correction changes runtime behaviour or crosses files outside the declared scope
Regression test failureCompare the base and proposed revisions; fix only when the intended rule is explicitThe pull request changes the rule, test data, or expected outcome
New security or authority pathDescribe the changed boundary and attach the relevant evidenceAlways require an accountable reviewer; never auto-approve
Migration or compatibility riskIdentify affected versions and the missing forward or rollback pathBlock until adoption order and recovery are decided
Unrelated existing defectCreate a separate finding linked to the evidenceDo not enlarge the current pull request merely because a nearby defect was found
The cleanest pull request is not the one with no warnings. It is the one whose remaining decisions are visible to the people authorised to make them.

Autonomy ladder

Widen evidence before widening authority

01 · decision

Observe

Summarise and classify reports without code access.

02 · decision

Recommend

Suggest priority, affected area, and reproduction steps.

03 · evidence

Prepare

Create a bounded change with tests, but never publish it.

04 · refused

Release narrowly

Only a defined, reversible fault class may cross the gate automatically.

The last step is not the target state for every team. Most value arrives while release authority remains human.

The ladder is a risk model, not a maturity contest. A team can gain most of the value at prepare while keeping every production release behind human authority.

Authority

Autonomy should follow fault classes, not model confidence

A model confidence score is not a production control. It describes a system's internal assessment, not the consequence of being wrong. A high-confidence change to tax calculation may deserve more scrutiny than a lower-confidence correction to an internal label. The useful unit of authority is therefore a named fault class with known blast radius, test evidence, reversibility, and an accountable owner.

An organisation might permit automatic release of a deterministic documentation correction or a dependency patch that passes a defined compatibility suite and can be rolled back without data change. It should not generalise that permission to “high-confidence bugs”. Each additional class is earned through observed performance and incident review. We examined the operational boundaries, safeguards, and interactive refusal queue for this model in Bugs fixed overnight, reviewed in the morning.

  • Start with read-only observation and make the evidence useful before permitting code preparation.
  • Require a reproducible failure or explicitly label the change as a hypothesis.
  • Constrain change size, affected paths, dependency scope, and permitted test commands.
  • Treat schema, identity, money, and irreversible external actions as separate high-consequence classes.
  • Revoke authority automatically when evidence is missing, telemetry is degraded, or rollback is unavailable.

Reproduction

A reproducible failure is a maintenance artifact, not a comment on a ticket

“Confirmed” is too weak a state for an automated maintenance path. A useful reproduction names the starting state, the action, the observed result, and the result that should have occurred. It identifies the product and release, controls time and external dependencies where possible, and fails for the same reason as the reported problem. Without that last condition, a passing fix may only silence a test that never represented the incident.

The best form is an automated test at the lowest boundary that still preserves the fault. A calculation error may need a unit test; a duplicate reservation may require a concurrency test against a real database; a provider-specific sign-in failure may need a contract fixture and a browser path. For an intermittent production fault, the first artifact may instead be a trace, a sanitised event sequence, or a deterministic replay. The format follows the failure rather than a universal testing preference.

Reproduction qualityWhat it establishesWhat remains unknown
Report onlyA user observed an unwanted outcomeStarting state, frequency, affected versions, and cause
Repeated manuallyThe outcome can be produced under known stepsWhether the repair will remain protected after release
Automated at the boundaryThe failure is repeatable and can become a regression gateWhether adjacent workflows or historic data also need repair
Production trace or replayThe real sequence and system state are representedWhether sensitive data is safe to retain and whether replay changes external systems

When reproduction is impossible, Bug Buster should preserve that uncertainty. A change can still be prepared as an instrumented hypothesis: add a narrow diagnostic, improve a refusal signal, or guard a suspected state transition. It should not be presented as a verified fix. Keeping “observed”, “reproduced”, and “explained” as separate states prevents the queue from turning confidence into fact merely because a plausible patch is available.

Prioritisation

Severity is a decision about consequence, not message volume

Ten identical reports may describe one contained inconvenience. One quiet reconciliation failure may corrupt every future statement. Prioritisation needs several dimensions: current user harm, data integrity, security or authority impact, number and identity of affected parties, persistence, detectability, workaround quality, and reversibility. Frequency matters, but it does not stand alone.

DimensionQuestionEscalating evidence
AuthorityCan a person see or do something they should not?Cross-tenant access, privilege persistence, missing revocation
IntegrityCan the system create an incorrect durable state?Financial mismatch, duplicate action, unreconciled record
ReachWho is affected now and on which versions?Growing cohort, shared dependency, no safe containment
RecoveryCan the outcome be reversed without losing legitimate work?Manual repair, external side effect, uncertain history
VisibilityWould ordinary monitoring reveal continued harm?Silent drift, successful responses, delayed discovery

Operations

Measure whether the maintenance system improves control

Closing more tickets can conceal worse maintenance. Useful operating signals include time to a reproducible example, the age of unbounded high-consequence reports, escaped regressions, reopened changes, rollback frequency, and the proportion of prepared fixes rejected because the intended behaviour was wrong. Those rejections are not wasted automation; they show the review boundary is doing work.

Every prepared change should retain its lineage: original observations, gathered diagnostics, reproduction, affected versions, code change, test evidence, approval, release, and post-release result. Incident review can then improve both the product and the common maintenance path. A recurring identity defect may become a shared contract test; a repeated pipeline escape may become a platform gate.

The system is learning when one product's failure becomes a guardrail for the rest of the ecosystem.

The counter-case

Some reports should never enter an automated repair path

Security disclosures, suspected fraud, legal holds, sensitive personnel matters, and active safety incidents need deliberately limited handling. They may require a protected intake, restricted evidence access, preservation rules, and coordination that ordinary triage must not automate or expose. The correct plugin behaviour is to recognise the class, capture the minimum, and transfer authority.

Small products with low change volume may also need only a disciplined checklist and existing delivery pipeline. Bug Buster earns its shared role when products repeat the same evidence-gathering and maintenance controls. It should not manufacture process where the coordination cost exceeds the risk it removes.

Continue through the system

Bugs fixed overnight, reviewed in the morning

How an in-house maintenance service works deferred defects on idle capacity with human review queues.

Read the article

The reusable unit is assurance

Why delivery speed depends on reusable evidence, controlled adoption, and product-specific gates.

Read the article

Share utilities without sharing business logic

How common capabilities can centralise policy and operations without absorbing the product's decisions.

Read the article

Let's talk about your challenge

If your organization is working with complex digital systems or exploring operational AI, we are always open to a conversation.